The Hidden Cost of AI in Legal Technology
Why Frontier Models Create New Risks
Legal has spent two years worrying about whether AI is secure, accurate and auditable. Harvey just exposed another risk: what happens when your entire technology stack depends on intelligence whose price you do not control?
Harvey had a remarkable problem this year.
According to Bloomberg , the $15.5 billion legal AI company began 2026 with gross margins of approximately 50 percent. Then Harvey updated its AI agents in March. Customer usage spiked, the agents consumed vastly more AI, and by June Harvey's gross margin had fallen to approximately negative 50 percent.[1]
That swing is hard to overstate. At a 50 percent gross margin, roughly fifty cents of direct cost supports each dollar of revenue. At negative 50 percent, the direct cost is roughly $1.50. In a matter of months, one of the largest and best-funded legal AI companies in the world found itself upside down on the basic economics of delivering its product.
Harvey has since pointed to investments in model routing, specialized infrastructure and open-weight models as ways to improve those economics. Its Tenet research specifically emphasizes reducing inference costs, and Harvey has published other work aimed at performing high-volume legal tasks at a fraction of the cost of frontier models.[2] Maybe those efforts have solved the problem Harvey faced this spring. The important fact for the rest of us is that the problem existed in the first place.
Harvey's product did not fail. Customers did not disappear. The AI became more capable, customers asked it to do more, and the amount of rented intelligence required to provide the service exploded. The economics changed underneath the product.
That should concern anyone building a law firm, legal department or technology business around frontier AI.
You do not control the price of intelligence you rent
Legal has spent enormous energy worrying about the right AI risks: confidentiality, privilege, security, hallucination, auditability and reliability. There is another dependency hiding underneath all of them.
If your technology sends essentially every meaningful task to OpenAI, Anthropic or another frontier provider, you do not control the unit economics of your own product.
You may control the user interface, workflow, prompts and customer relationship. You may have built a terrific application. But if the intelligence that makes it work is rented, somebody else controls the economics of the most important input into your product.
Right now, frontier intelligence is being produced at extraordinary cost. The Financial Times reported this month that OpenAI expects almost $280 billion of negative free cash flow between 2026 and 2030, despite projecting enormous revenue growth. OpenAI expects total spending through the end of the decade of approximately $856 billion.[3]
That cannot remain economically disconnected from what customers ultimately pay forever.
Maybe chips get radically cheaper. Maybe models become dramatically more efficient. Maybe competition continues driving down the published price of individual tokens. But unless efficiency gains outrun both exploding demand and the increasing amount of computation required to perform more sophisticated work, the frontier providers eventually have to extract more economic value from the intelligence they provide.
That does not require an email saying, "API prices are doubling."
It can mean premium models, specialized models, agent charges, long-context premiums, faster-processing tiers, capacity pricing or simply workflows that consume vastly more inference than their predecessors. OpenAI already charges differently for different levels and modes of intelligence.[4]
Harvey demonstrated the other side of the equation. The price of the individual input did not need to explode. The quantity of intelligence the product consumed did.
That distinction matters. The question is not merely, "What happens if OpenAI doubles its prices?" It is, "What happens if delivering the level of intelligence my customers expect suddenly costs two, five or twenty times more?"
A cheaper token is not much comfort if tomorrow's workflow needs twenty times as many of them.
OpenAI is not just underneath legal AI anymore
This risk becomes more interesting as the frontier companies move directly into legal.
On September 17, OpenAI introduced Astra for Law, a version of GPT-6 Astra specifically tailored for professional legal work. OpenAI says it can research legal questions based on client facts, develop arguments and deal terms, analyze how contractual provisions allocate risk, and search a dedicated index of U.S. case law, statutes, regulations and other authorities. It is initially being offered to selected U.S. law firms and is also expected in the API.[5]
OpenAI is no longer simply providing infrastructure to companies building legal AI. It is increasingly selling legal intelligence itself.
And legal intelligence already commands a premium.
OpenAI's current enterprise token rate card prices ordinary GPT-6 Astra at $10 per million input tokens and $50 per million output tokens. GPT-6 Astra Law is $12.50 and $62.50 respectively. That is already a 25 percent premium for the legal version of the same model family.[6]
Will that premium be 25 percent in three years? Ten percent? One hundred percent? I have no idea.
And neither does the law firm reorganizing its work around it.
That is the risk.
If your AI strategy is built directly around OpenAI, or around a legaltech application whose architecture is fundamentally a chain of calls to OpenAI, Anthropic or another frontier provider, a critical part of your future cost structure is controlled outside your organization. If the provider changes its pricing, introduces a materially better premium model, changes its product packaging, or if the next generation of agentic work simply requires much more inference, you inherit those economics.
And legal is actively increasing that dependency. We are telling lawyers to research with AI, draft with AI, review with AI, analyze with AI and build agents around AI. Clients are being told those tools should make legal work faster and less expensive. Staffing and billing expectations will inevitably begin to reflect those efficiencies.
Once that happens, "we'll stop using it if it gets too expensive" stops being a serious answer.
If the underlying economics change, the frontier provider passes the cost to the legaltech vendor. Or the vendor's own usage costs simply explode, as reportedly happened to Harvey. The vendor can absorb the cost, restrict the product, substitute a less capable model or raise prices. Eventually that cost reaches the law firm, legal department or client.
That is the trickle-down economics of rented intelligence.
The same architecture creates the security problem and the cost problem
Then there is the app someone in Legal Ops vibe-coded last Tuesday.
I wrote recently about the security problems surrounding these applications. A useful internal tool can now be built in an afternoon by someone who is not an engineer. But that does not mean anyone has reviewed its authentication, mapped where its data travels, verified its subprocessors, examined how credentials are stored, implemented appropriate logging or determined whether its output can later be reconstructed and defended.
The economic problem comes from exactly the same architectural shortcut.
Ask what the application itself actually knows.
Often, very little. It knows how to ask an LLM.
An invoice goes in, so the LLM interprets it. A contract comes in, so the LLM extracts its terms. Something needs classification, so the LLM classifies it. Analysis is required, so the LLM analyzes it. A more sophisticated workflow is needed, so an agent makes a series of LLM calls, invokes tools, examines the results and calls the model again.
The same design that creates questions about security, auditability and repeatability also means there is usually no independent economic architecture underneath the application. No proprietary schema. No deterministic analytical engine. No specialized model infrastructure. No body of domain intelligence capable of doing some of the work without renting intelligence from someone else.
If the underlying model becomes more expensive, the application becomes more expensive. If tomorrow's agent requires ten times more inference, it consumes ten times more inference. There is nowhere else for the work to go.
Harvey at least has hundreds of millions of dollars and teams of engineers with which to attack that problem. The application somebody built in a browser last week does not.
The companies that built something before the LLM may have an advantage
There is a strange irony in all of this. The rush toward "AI-native" legal technology has sometimes treated anything built before generative AI as obsolete baggage.
I think the opposite may prove true.
Some of the legal technology companies best positioned for this transition may be precisely the ones that spent decades building proprietary domain intelligence before ChatGPT existed.
Consider Thomson Reuters Westlaw. Its West Key Number System is a proprietary taxonomy covering more than 140,000 granular legal categories. KeyCite contains more than 1.4 billion connections linking legal authorities and that taxonomy. Thomson Reuters says it performs more than 1.6 million editorial enhancements each year, creating a proprietary layer of structured legal intelligence around the underlying law.[7]
Or consider LexisNexis. Lexis+ with Protégé now incorporates generative and agentic AI, including Anthropic capabilities, but those systems are grounded in LexisNexis's existing repository of roughly 200 billion legal documents, with millions added daily, together with Shepard's and its linked and structured legal content.[8]
Neither company is avoiding generative AI. That is not the point.
They have something underneath it.
An LLM can enhance Westlaw without requiring Westlaw to ask an LLM to rediscover the structure of American law from scratch on every query. Lexis can integrate Claude without making Claude synonymous with the entirety of Lexis's legal intelligence.
That does not eliminate their exposure to inference costs. It gives them options that a wrapper does not have.
The more work that can be performed using proprietary taxonomy, retrieval, citation networks, structured data, deterministic logic, specialized models, and other domain infrastructure, the less completely the product's economics are dictated by frontier inference.
The old moat may turn out to be the new moat.
That is why the architecture underneath Legal Decoder matters
Legal Decoder sits in a different part of the legal market, but the architectural principle is the same.
We began building technology to understand legal billing long before ChatGPT existed. Over more than a decade, we developed a proprietary machine-learning schema describing what lawyers actually do and how that work appears in legal invoices. We built specialized classifications, mappings, analytical structures and metrics around that domain.
Generative AI did not make that work obsolete. It made it more valuable.
We do not need to rent frontier intelligence to rediscover everything our technology already knows. If a billing entry can be mapped against an existing proprietary structure, there is no reason to ask the world's most expensive general-purpose intelligence to recreate that structure every time an invoice arrives. If a metric can be calculated deterministically, there is no reason to calculate it probabilistically. Repetitive high-volume workloads can be handled with specialized or self-hosted systems where appropriate.
Then, where a frontier model does something extraordinary, use it.
This is not an argument for less AI. It is an argument for using the right intelligence for the right job.
That architecture has security advantages because less information needs to leave systems we control. It has reliability and auditability advantages because deterministic work can remain deterministic. And it creates an economic escape valve. If frontier intelligence becomes materially more expensive, every component of the product does not automatically become materially more expensive with it.
An application built entirely around an external LLM cannot create that escape valve after the fact without building something underneath the LLM.
And the frontier model provider has an even more fundamental limitation. OpenAI can make Astra for Law extraordinarily capable. But OpenAI cannot place some independent layer of intelligence beneath OpenAI to avoid calling OpenAI.
It is the model.
If your strategy is that OpenAI performs the legal work, the economics of performing that work remain tied to OpenAI's economics.
That does not make OpenAI bad. It makes OpenAI a critical supplier. The difference is that this supplier may eventually be embedded in an enormous percentage of the substantive work your lawyers perform.
Harvey is the warning
Harvey has said it is addressing its inference economics through open-weight models, post-training, routing and more efficient infrastructure.[2] But whether Harvey has solved Harvey's current problem is not really the important question.
The important fact is that one of the most sophisticated and heavily funded legal AI companies in the world reportedly went from roughly positive 50 percent gross margins to roughly negative 50 percent margins within months because the underlying consumption economics moved against it.
That should not reassure legal buyers. It should make them start asking harder questions.
What percentage of this product's work requires frontier inference? Which functions can operate without it? What proprietary intelligence exists underneath the model? Which outputs are deterministic? Which workloads can run on specialized or self-hosted models? What happens to my price if the model provider's effective cost doubles? What happens if the next version of your agent uses ten times as much compute?
And there is one question that cuts through almost all of it:
If the LLM disappeared tomorrow, what intellectual property would still be left?
If the answer is a user interface, some prompts, workflows, and API integrations, understand what you are buying. You are not simply buying software. You are buying somebody else's access to rented intelligence.
The bill eventually reaches the customer
It is possible that frontier AI becomes so inexpensive that none of this matters. Hardware could improve faster than usage grows. Open models could commoditize large portions of inference. Competition could keep prices falling.
But "the critical input into our business will keep getting cheaper forever" is not a cost-management strategy. It is a bet.
OpenAI is projecting hundreds of billions of dollars of negative free cash flow while simultaneously moving further into specialized professional markets. Agentic systems are being asked to perform increasingly complicated and compute-intensive work. Legaltech companies are building larger portions of their products around that intelligence while law firms and legal departments restructure their own work to depend on it.
Harvey just showed how quickly those economics can move in the wrong direction.
By the time the bill arrives, your lawyers depend on the technology. Your clients expect the efficiency.
At that point, the question will not be whether you still want AI. It will be whether there is anything underneath your AI that gives you another option.
The durable legal technology companies will not simply be the ones that make the greatest number of calls to the smartest model available. They will be the ones that own enough proprietary intelligence to know when a frontier model is necessary, when it is not, and what to do when the economics of renting that intelligence change.
The most important AI model in your stack may be the one you don't call.
References
[1] Rebecca Torrence and Natasha Mascarenhas, “OpenAI, Anthropic Costs Push More Startups to Build Off Cheaper Open Models,” Bloomberg, September 21, 2026. Bloomberg reported that Harvey's customer usage spiked following a March agent update and its gross margin fell from approximately 50 percent at the beginning of 2026 to negative 50 percent by June.
[2] Harvey, “Harvey Tenet Research Preview,” August 20, 2026, and related Harvey research on specialized models. Harvey describes using post-training and open-weight models to improve cost efficiency and reduce tokens consumed at inference time. Its separate work on high-volume review workloads reports specialized models performing at a fraction of frontier-model cost.
[3] Financial Times, “OpenAI expects to burn $280bn by 2030,” September 19, 2026. The FT reported from a company presentation that OpenAI projects nearly $280 billion in negative free cash flow between 2026 and 2030 and approximately $856 billion of spending through the end of the decade.
[4] OpenAI API pricing, accessed September 24, 2026. OpenAI currently differentiates pricing by model, context length and processing tier, with additional charges or premiums for certain features and modes.
[5] OpenAI, “Astra for Law,” September 2026. OpenAI describes Astra for Law as GPT-6 Astra tailored for professional legal work, supported by a dedicated Legal Search Index and initially offered to selected U.S. law firms.
[6] OpenAI, “ChatGPT Rate Card: Enterprise Token-Based Pricing,” accessed September 24, 2026. The current rate card lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens, compared with $12.50 and $62.50 for GPT-6 Astra Law.
[7] Thomson Reuters investor materials, 2026. Thomson Reuters describes the West Key Number System as a proprietary taxonomy covering more than 140,000 precise legal categories and KeyCite as a proprietary citation network containing more than 1.4 billion connections.
[8] LexisNexis, “LexisNexis Expands Lexis+ with Protégé by Integrating Anthropic's Claude Legal Plugin Suite,” May 13, 2026. LexisNexis says its generative AI capabilities are grounded in a repository of approximately 200 billion legal documents containing Shepardized and linked legal content.