Jev vs Frontier LLMs
Primer
Every day, millions of datapoints on titles, talent, and companies flowed into Netflix’s Entertainment Knowledge Graph from hundreds of sources, and each one had to be matched to the right entity in the graph. Up to this point, most of the graph was constructed using a suite of ML models, supported by data labelers for edge cases and LLMs for a minority of sources. Both of these supporting solutions had been cost prohibitive to deploy at much larger scale to handle trickier edge cases. Recently, I reran a small version of that problem on a new kind of large language model, and it matched Claude Opus 5.5’s accuracy for about 1% of the cost.
Background
I spent seven years leading product for that graph, which Netflix deployed across the company to inform content licensing, subscriber pricing, competitive positioning, and personalization on the service:
Entertainment knowledge graph. Source
A team of multidisciplinary scientists, engineers, and labelers worked on ensuring the graph had sufficient coverage, freshness, and accuracy across the suite of applications. One of the largest challenges was accurately representing the state of the entertainment world through a process called Entity Resolution: the process of identifying, matching, and linking different records or references that refer to the same real-world entity across heterogeneous and evolving sources.
The original set of algorithms that handled Entity Resolution, specifically the matching step, performed well with source entities that had sufficient, well structured metadata. However, it performed worse with source entities that arrived with little or loosely structured metadata. The ML team that owned matching experimented with attempts to onboard these sources quickly and accurately, but because each source had unique differences that were difficult to generalize, less than desired progress was made to incorporate them into the graph. When critical, many of the low confidence source entities were onboarded with the help of third party data labelers to aid with the matching process.
Enter: LLMs
Human experts were a decent stop gap–but an expensive one, especially when local language expertise and domain knowledge was required. Large language models–equipped with web search tooling–turned out to be the next solution. After much experimentation, these models were able to match these entities at a fraction of the cost of human labelers for some use cases.
While adding LLMs in the loop helped some last mile matching problems, they were not a complete substitute for the ML-based matching process. Each day, millions of entities needed to be evaluated for matching, and since the cost to deploy the LLM-based matching solution with high enough accuracy was prohibitively high, this LLM process could only be effectively deployed where the cost was justified at smaller scale.
Enter: TypeSafe AI’s Jev
On Sept 15, 2026, TypeSafe AI released Jev, which it calls the first “System One” model. Instead of generating language, Jev answers a fixed set of typed questions (choose one option, give a score, or yes/no) with structured values that software can use directly. That is exactly the shape of a matching decision.
There is a LOT of AI noise out there. Filtering out the breakthroughs from the noise has become a skill itself. When I heard about the promise of Jev, I was excited–so much of my scaled work with LLMs had been using structured outputs via the APIs. So I decided to figure out whether the product lived up to the marketing promise.
Jev: Decision model for classification tasks
To test it, I recreated a small version of the Entertainment Knowledge Graph matching problem. I created an evaluation dataset by sourcing a few hundred popular titles from a Spanish top tv ranking website and manually matched the correct Wikipedia page for each program (or supplied none if no page exists), as a human annotator would do. Next, I sourced web context search data via Serper.dev so that each model would be making decisions based on the same internet context. Finally, I developed scripts and provided them with the matching task and the search results. I selected a leading frontier model (Opus 5.5), more cost effective open source model (Qwen 3.8 Max + Flash), and lastly the challenger Jev to evaluate accuracy and cost per task performance.
The Results
In success, each model would reach 95% accuracy or higher. On the first pass, Jev lagged far behind at 52%. After I split Jev’s decision into two questions (which candidate is best, and is any candidate right at all), Jev reached 96%, the same accuracy as Opus 5.5 on the same candidates. Qwen 3.8 Max scored 96%.
The big surprise, however, was cost. Jev was about 100x cheaper than Opus 5.5, 70x cheaper than Qwen 3.8 Max, and just over 5x cheaper than Qwen 3.8 Flash, which also scored lower than the threshold (89%).
Jev: 100x cheaper than Opus 5.5 and 5x cheaper than Qwen 3.8 Flash
A two-orders-of-magnitude cost reduction at comparable accuracy is a breakthrough. At Netflix scale, millions of matches a day, that is the difference between a budget line item and a rounding error. Assuming that this generalizes to many other classification tasks, I can see System One models like this becoming the default for high-volume classification. This was only a small test for one matching task against a small sample set. But the scale of improvement is hard to ignore.
⬅️ back to writing |












