Hero Background

Next-Gen App & Browser Testing Cloud

Trusted by 2 Mn+ QAs & Devs to accelerate their release cycles

Next-Gen App & Browser Testing Cloud
Testμ

How Startups Are Rethinking Value and Monetization [Testμ 2026]

Vibhor Rastogi of Citi Ventures on why one license can run a thousand agents, what that does to unit economics, and why outcome pricing still has no referee.

Published on:

Buy one software license. Point a hundred agents at it. Or a thousand.

The vendor’s token bill scales with the agents. The vendor’s revenue stays fixed at the seat. That gap is what this session is about.

At Testμ Conf 2026, Vibhor Rastogi, Managing Director at Citi Ventures, worked through what replaces per-seat pricing when agents are the users.

Youtube thumbnail

If you couldn’t catch all the sessions live, you can access the recordings at your convenience by visiting the TestMu AI YouTube Channel.

TL;DR

Seat-based pricing with capped consumption is a software pricing model that charges per user license and limits how much usage each license entitles the customer to. It exists because agents broke the seat: one license can drive a thousand agents, so unlimited usage scales the vendor’s token cost while revenue stays flat.

  • Is per-seat pricing still viable for agent products? - No, because one license can drive an unlimited number of agents. Vibhor Rastogi’s point is that a customer can buy a seat and run 100 or 1,000 agents against it, so if the vendor sold unlimited tokens at a fixed per-seat price, the token cost scales while the revenue does not.
  • What is replacing the per-seat model for agent products? - Seat pricing with a consumption cap, at least for now. Vibhor Rastogi says the pattern is already visible in coding agents and chat products, which advertise a per-seat price and attach throttling on weekly usage, daily usage and which model tier the seat may call.
  • Is outcome-based pricing ready for agents today? - No. Nobody has settled who defines and who measures the outcome. His example: a voice agent reports it resolved the call, the customer calls back five minutes later, and it is unclear whether the vendor gets paid.
  • Who is at fault when an agent fails on bad data? - Nobody has settled it. If a voice agent queried a knowledge base that did not hold the right data, Vibhor Rastogi asks whether the vendor is penalised, the customer who maintains that knowledge base is, or the vendor should take the knowledge base over too.
  • Has outcome-based pricing been tried before AI? - Yes. Vibhor Rastogi points to GE selling industrial equipment and software on outcome terms, where customers paid if it cut operating cost or kept more planes flying. His lesson is that the contracts became complicated over measuring uptime and attributing results to the vendor.
  • Why did seat-based pricing win in the SaaS era? - Because the seat killed shelfware. Under pre-SaaS licensing, buyers purchased far more capacity than they needed and left it unused, while per-seat pricing meant paying for what you used, scaling up with headcount and down in a downturn.
  • How many more tokens does an agentic product consume than a chatbot? - Vibhor Rastogi’s figure is 10 to 50 times more, offered without a source. His explanation is that every agent turn ships the model a large context and multi-turn conversation compounds it.
  • What is output maxing versus token maxing? - Token maxing aligns a vendor with its own consumption rather than the customer’s value, and Vibhor Rastogi argues for output maxing instead. Once the cost of delivery exceeds the value received, the customer churns, switches provider, or builds it themselves.
  • Should agents be allowed to spend money autonomously? - No, not without a ceiling. Vibhor Rastogi expects the same construct as consumption caps to apply to payments, so agents get autonomy but never uncapped payment authority. He runs a personal shopping agent capped at a $25 threshold.
  • What should an AI startup build that enterprises will buy? - Three categories: eval infrastructure, meaning the tooling rather than the data feeding it; agentic and AI security; and the knowledge graph and semantic ontology layer. He cites acquisitions in observability and AI security as evidence enterprises would rather buy than build.
  • What should enterprises keep proprietary in their AI stack? - Context, which Vibhor Rastogi calls the proprietary crown jewels, along with any logic that enriches it. He extends the rule to the harness and evals, with a blunt buying implication: do not buy a product where you do not own the harness.
  • Is one layer of the AI stack capturing most of the value? - No. Vibhor Rastogi refuses to pick a layer, arguing value is distributed unusually evenly across chips, foundation models, applications and infrastructure because enterprise tech spend is rising across all four.

The Investor Vantage Point

Vibhor Rastogi says he has spent 15 years in venture capital, starting at Intel Capital, then a private equity fund, and almost six years at Citi, with essentially the whole career on growth-stage software.

His entry point is post product-market-fit companies with a couple of million in ARR that have raised friends-and-family and seed money and are raising their first institutional round, after which Citi Ventures stays invested through to a liquidity event.

The disclosure that matters for everything below: the portfolio names he claims personal involvement in include Glean, ClickHouse, Galileo and Lexion, and he says Citi Ventures was an investor in Lakera. Those companies then supply most of the session’s supporting examples.

The group runs two practices, enterprise software covering cloud, cyber security, DevOps and AI and data infrastructure, and fintech covering payments, lending, wealth and commerce.

He dates his own AI investing to about a decade ago, starting with DataRobot, and notes the private equity fund he was at focused exclusively on artificial intelligence around 2018 and 2019.

Why Did Coding Break It Open?

He frames AI as a field with a documented boom-and-winter cycle going back to the 1960s, and argues that compute, algorithms and data converging is what made this cycle enterprise-grade for the first time.

His sharpest distinction is that the deep learning and computer vision breakthrough mostly benefited consumer internet companies and had little enterprise application, because most enterprises are not working on image data.

Large language models changed that because a lot of enterprise data is unstructured, and so is coding. He puts unstructured data at 80% of what enterprises hold, a widely circulated figure he offers without a source.

He calls the coding agent the archetype example of AI going mainstream, and reaches for revenue growth at coding companies as evidence. Those figures come from memory, without sources, and he uses revenue and ARR interchangeably, so they are best read as an impression of pace rather than as numbers.

He also makes a very large claim about a foundation model lab reaching a revenue scale he says is unprecedented in software history, most of it from coding. It is his own projection, no source is given, and the figures are extraordinary enough on their face that this recap does not reproduce them.

The structural point survives the numbers. Coding taught the industry the deployment pattern, meaning what skills an agent needs, what memory is, what a harness is, how evaluation works and how sandboxes work, and once those patterns lock down, voice, sales, marketing, customer service, HR, legal and vertical agents follow.

What Does Deployment Take?

Asked what the market underestimates, he answers that it is working through everything involved in deploying agents inside an enterprise, and reaches for PolyAI, a portfolio company, as the worked example.

He describes voice agents deployed by restaurant customers running around the clock with natural multi-turn conversation, accent handling, empathy, latency control and noise reduction, and notes they disclose to callers that they are talking to a bot rather than leaving it ambiguous.

The visible agent is the easy part, in his telling. First you solve how the agent is fed context: where it learns how your company does business, how your human agents respond, what knowledge bases exist and what gets escalated to which support tier.

Then come skills and tools, because the agent needs to look information up and reason through whatever a caller throws at it. Only then does the enterprise have a working deployment.

Security is his third pillar, framed as something humans handle instinctively and agents do not, which he turns into a question about guardrails, security policy and posture: how do you make sure AI does not break AI.

He references a recent incident in which a coding agent was reward hacked and began downloading files from a public model repository. He gives no date or source and the model version is not intelligible in the recording, so no product is named here.

Enterprises Lag The Tech

The host puts the pessimist’s case: agents have improved coding velocity, that has not translated into release velocity, so expectations may be running ahead of delivery.

Rastogi rejects the framing and inverts it, arguing enterprises are perhaps more conservative than they need to be and the technology is running ahead of where enterprises are.

He uses Galileo, another portfolio company, as evidence of the testing burden, describing production evals that checked whether a consumer chatbot’s responses were unbiased, non-toxic, truthful and grounded.

His diagnosis of the delay is fear rather than capability. Enterprises spend a fair amount of time testing these systems before they will put them out, because everybody is worried about the unintended consequence.

As counter-evidence he offers his own portfolio, saying 70 to 80% of their code is now written by AI at companies with excellent engineers who recognise AI can write better code than they can. That is a second-hand, self-reported number about unnamed companies he has a financial interest in.

He does concede a limit, invoking the jagged frontier: some use cases where AI beats 95% of humans, maybe 99%, and others where a middle schooler does better, naming legal AI as a domain where accuracy really matters. The hedge in that figure is his own.

Note

Note: Know what a run costs before you price it - and before you ship it. Try TestMu AI now!

Own Or Rent Intelligence

On build versus buy he borrows a framing he attributes to Satya Nadella, asking whether you want to own intelligence or rent it, and sharpens it: if your data, your model and the surrounding infrastructure are all owned by third parties, what is your true competitive advantage?

He lays out the opposite extreme too, building the model, fine-tuning it and owning every component of the harness, and refuses both. Neither extreme is right, and the answer is a middle ground.

On foundation models he is categorical that you should not try to build one, because the capital raised by the major labs is beyond what very few companies can match.

What remains available is fine-tuning, reinforcement learning and supervised fine-tuning on top of a foundation model, plus cost optimisation through using prior-generation models for routine tasks and reasoning models for harder ones, which is why he predicts model routing becomes an important capability.

Context is the part he says must never be rented, calling it the proprietary crown jewels along with any intelligence that augments it. His example is using AI to score and enrich proprietary customer data, where the scoring and enrichment logic should stay owned by the company.

He extends the rule to the harness and evals, the loop where the model produces an output, you tweak it and feed it back. The buying implication is blunt: do not buy a product where you do not own the harness, and be wary of third-party products that own context, model and harness at once.

Evals, Security, Graphs

Pressed on what a startup should build that enterprises will actually buy, his first answer is eval infrastructure, with the caveat that the startup sells the infrastructure and not the data feeding the eval.

As market proof he points to recent acquisitions in AI observability and AI security. The acquirers and amounts are either garbled in the recording or unverified, so they are not reproduced here.

His second category is agentic security, or AI security generally, on the reasoning that enterprises will not want to build it even though it makes the end-to-end workflow far more secure.

Third is the knowledge graph and semantic ontology layer, the technology the context is fed into, which he says companies are not going to build themselves.

He points to two portfolio companies as live examples, one commercialising a context layer plus an agent-building framework for large enterprises, and one selling agentic data management on a high-performance real-time database.

The division of labour he sketches is that startups build these infrastructure components while enterprises consume them to build their own context layer and harness, and buy the foundation model.

Why Does Per-Seat Break?

The host sets up the pricing segment by describing the old model, humans with a login and everything funnelled down to a seat count, and asks what replaces it. That framing is his rather than the guest’s, and the published chapter list blurs the two.

Rastogi’s answer is a unit-economics argument rather than a philosophical one. He could buy a license and run 100 agents with it, or a thousand.

The startup-side consequence he states plainly: if you were selling unlimited tokens at a fixed per-seat price and your customer converted that seat into a thousand agents, you would have fairly bad unit economics.

This is not theoretical advice on his part. He says Citi Ventures has its portfolio companies think through that dynamic very clearly before setting a price.

He is careful to explain why the seat model won in the first place. Before SaaS, enterprise licensing agreements meant buying well over the capacity you needed and leaving it on the shelf, and per-seat pricing was the innovation that ended shelfware, with Salesforce as the example: ten sales reps, ten seats, more when you grow, fewer when you shrink.

Test across 3000+ browser and OS environments with TestMu AI

The Outcome-Pricing Trap

His objection to outcome-based pricing is not the principle but the referee: who defines the outcome, and who measures it?

Comma

His second case splits the blame differently. The voice agent queried the knowledge base and the knowledge base did not hold the right data, so is the vendor penalised, or the company that was supposed to maintain that knowledge base, or does the vendor now take over the knowledge base as well?

He notes outcome pricing long predates AI, giving GE as the precedent: industrial equipment sold with software where the customer paid if it genuinely reduced operating cost, or for jet engines, kept more planes in the air. He recounts it from memory, without dates or contract specifics.

The lesson he draws is contractual rather than technical. Those contracts became very complicated over how to measure uptime and downtime, and how much of the result was truly attributable to the vendor.

He still treats outcome pricing as the destination rather than a dead end, describing the shift as a phased journey from seat-based to outcome-based, and saying that if the industry can work out how to do it, that is the best result for customer and vendor alike.

Seat Caps And Payment Caps

ModelHis verdict
Per-seat, unlimited usageBreaks, because one seat can drive a thousand agents against the vendor’s token cost
Pure outcome-basedThe destination, but unresolved: no agreement on who defines or measures the outcome
Seat plus capped consumptionWhere he expects the market to settle in the interim

He says the capped model is already visible. Coding agents and chat products advertise a per-seat price that comes with throttling gears: weekly usage, daily usage, and the class of model you are allowed to use.

Asked whether regulated financial services makes outcome definition harder, he declines the distinction, saying there is no nuance between a regulated and an unregulated enterprise because it is a general financial question for the CFO. He reframes pricing as risk allocation: what the buyer keeps and what they push onto the vendor. He never supplies a method for quantifying outcomes, which was the question.

On enterprises wanting a price signed in January to hold until December, his answer is that the cap itself is the mechanism. Spend per license will not exceed the seat price plus the consumption cap, and the customer can raise the cap later in the year if it runs out early.

When the host raises agentic commerce, a coding agent with access to your credit card buying the tool it prefers, he answers that he does not think so, and extends the same construct to money: guardrails, thresholds and payment caps, so agents get autonomy but never authority for uncapped payments.

His supporting example is personal rather than portfolio, and it is the one genuinely first-hand artefact in the session. He runs a shopping agent that buys on his behalf under a $25 threshold, and compares the coming pattern to corporate signature authority, where some approvals go to the CFO, some to the CEO and some to the board.

Token Math, Output Maxing

Asked whether getting pricing wrong at the start is fatal, he hedges twice and never delivers a verdict, saying he does not know if it is a death sentence and later that he would not call it detrimental. What he will say is that a founder who has not thought it through is in a bad position.

His number for why: an agentic startup’s models will consume 10 to 50 times more tokens than a simple chat-based application, because every turn sends the foundation model a large context and multi-turn conversation compounds it. The figure comes without a source.

He makes mispricing concrete with an explicitly hypothetical example: sell at conventional SaaS pricing of $50 or $100 a month while your own cost of delivery runs to tens of thousands a month, and you lose money on every customer you sign.

His mitigation is a discipline rather than a pricing model. Know how customers will actually use the product, hold a reasonable estimate of token consumption, and keep contractual flexibility to change pricing.

Comma

His voice-agent arithmetic is the clearest version of output maxing. A human-answered contact centre call costs six to seven dollars, so an agent delivering it for fifty cents to a dollar saves the customer several dollars a call, and if the agent instead costs the customer ten dollars, no economic value was delivered. The figures are round numbers offered without a source, and the saving he states does not track exactly against his own range.

Where Does Value Accrue?

Asked whether the money will be made at the application or infrastructure layer, he refuses to pick, describing a fairly good distribution of value between all the different parts of the value chain.

On chips he lists GPUs alongside the major cloud providers’ own silicon and a set of independent chip companies working on low-cost inferencing. He sizes that layer at several hundred billion dollars of revenue cumulatively, an aggregate given without a source.

Foundation model companies are his second growth engine, and the application layer his third, populated with coding, voice AI and sales AI companies.

The infrastructure layer holds data management, context, evals and observability, all of which he says are also doing well.

His conclusion is that no single layer is capturing disproportionate value, because enterprises keep shifting more of their technology spend toward this and it benefits everybody in the chain.

Subsidy And Open Weights

An audience member asked whether AI startups are genuinely capital efficient or whether inference and infrastructure costs are flattering the economics. The answer that follows comes from an investor being asked whether investors are inflating the numbers.

He names the historical version of the concern, the ride-hailing and food delivery era where venture money subsidised consumer prices below the cost to deliver.

His verdict is that this is largely not happening in AI, while conceding there are certain sectors where pricing may be lower than it needs to be. His argument is that cost pressure arrived so fast that cost optimisation started very early.

He casts open-weight models as a direct response to that pressure, saying the cost arbitrage is through the roof, with near state-of-the-art models available for well under a dollar per million tokens and costs falling by an order of magnitude year on year. All of those figures are unsourced and were current as of the recording.

Asked for a two-year prediction, he points to acquisitions of three- and four-year-old AI companies already happening and expects large corporates to keep buying. The specific acquirers and amounts he cites are garbled or implausible as transcribed, so they are not reproduced here.

He closes on a widening frontier, referring to a biotech result where foundation models helped predict new vaccines, therapies and cancer drugs, and to a robotics company going public, before listing autonomous cars, physical robotics, biotech, materials, space, quantum, data centers and software agents as all on an upward trajectory. No paper, institution or company is named clearly enough to cite.

This session was part of Testμ Conf 2026, which ran across three days of sessions on agentic engineering and quality. Registrations for the next edition are already open on the Testμ Conference 2027 page.

Author

...

TestMu AI

Blogs: 232

  • Twitter
  • Linkedin

TestMu AI is World's First Full Stack AI Agentic Quality Engineering platform that empowers teams to test intelligently, smarter, and ship faster. Built for scale, it offers a full-stack testing cloud with 10K+ real devices and 3,000+ browsers. With AI-native test management, MCP servers, and agent-based automation, TestMu AI supports Selenium, Appium, Playwright, and all major frameworks. AI Agents like HyperExecute and KaneAI bring the power of AI and cloud into your software testing workflow, enabling seamless automation testing with 120+ integrations. TestMu AI Agents accelerate your testing throughout the entire SDLC, from test planning and authoring to automation, infrastructure, execution, RCA, and reporting.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests