Harvey, the legal AI startup valued at $15.6 billion, saw its gross margins collapse from roughly 50% at the start of the year to negative 50% by June 2026, as Bloomberg reported.
- This guide provides comprehensive, actionable information
- Consider your specific workflow needs when evaluating options
- Explore our curated LLMs tools for specific recommendations
What actually happened
Harvey, the legal AI startup valued at $15.6 billion, saw its gross margins collapse from roughly 50% at the start of the year to negative 50% by June 2026, as Bloomberg reported. The cause was straightforward: token usage under OpenAI and Anthropic's usage-based enterprise pricing rose twentyfold as customers used the product more, and the bill scaled with it. Running at negative margins on your core product is not a sustainable position for a company of that size, whatever the growth numbers look like.
Harvey's fix was to stop renting frontier intelligence and build its own. In August 2026 it launched Tenet, an in-house model post-trained on top of Moonshot AI's open-weight Kimi K3. Harvey says Tenet runs at less than a quarter of the cost of leading foundation models while holding strong performance on its legal benchmarks, and margins turned positive again after the switch.
Harvey is not the only one
Bloomberg's reporting names other startups making the same move: Abridge, Decagon and Ramp are each building on cheaper open-weight models rather than staying fully dependent on closed frontier APIs, backed by investors including Sequoia Capital and General Catalyst. The pattern is the same in each case: a company with real usage volume finds that per-token enterprise pricing scales worse than its own revenue, and that training or post-training an open-weight base now costs less than continuing to rent.
Why this is economically viable now, and was not before
This move only works if the open-weight base is actually good enough, and that gap has been closing fast. By the same reporting, Moonshot AI's open-weight Kimi K3 has narrowed the performance gap to Anthropic's closed Fable 5 to about four months, at roughly a fifth of the cost. A year ago, the capability gap between an open-weight base and a frontier closed model was wide enough that most companies had no real alternative to paying whatever the API charged. That is no longer true for workloads, like Harvey's legal analysis, that are well-scoped enough to post-train a smaller open base against.
The broader effect Bloomberg points to is pressure on OpenAI and Anthropic's own margins, and a shift in where AI's actual profit lands, potentially toward chips and infrastructure rather than the model layer itself, as more of the companies that used to be their biggest customers become their competitors on the model layer.
Who should actually care
For what open-weight models still cannot do, and where this strategy runs into a wall, see open models: what they cannot do. For the build-vs-rent decision itself, see AI infrastructure: run, host, or rent. For the model at the center of this, see Kimi K3 and every LLM in the directory.
Explore curated tools related to this guide: