🔍 Read the full analysis: 24 Strategies For Using Jev In AI Decision Models on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer says he has mapped 24 uses for Jev in AI decision models, including three already running in his publishing operation and 12 he rates as strong fits. His recommendations are based on volume, narrow questions, the cost of errors and evidence that current rules fail; results cited in the article come from his own measurements.
Thorsten Meyer has mapped 24 uses for Jev, a tool that returns typed answers to narrow questions so software can route routine decisions. In a September 29 article, he says three uses are live in his publishing operation, 12 more meet his criteria for strong fits, seven need measurement first and two are poor fits.
Meyer describes Jev as a system that takes text or JSON plus typed questions and returns answers that code can act on. Its answer types include a yes-or-no probability, a choice with probabilities and confidence, or a score on ordered levels. He says a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. Jev does not write or summarize the material, according to his description.
Three of the listed uses are already operating in Meyer’s publishing workflow: checking whether stories fit a site, detecting non-English article text, and classifying headlines as a fallback when a primary language model fails. Meyer reports that a scan of 78,889 articles cost $2.01 and identified 1,576 non-English items, of which 1,553 were fixed. He also reports roughly 10,000 story-and-site relevance judgments over three days, with 22% clearly on-topic, and 89% agreement with a frontier model for the fallback classifier.
The proposed publishing uses include checking disclosures, moderating comments, assessing headline quality and detecting stories with too few verifiable facts. Meyer labels the thin-source detector, product matching in roundups and headline checks “measure first,” because he says the failure rate of existing methods has not yet been established. He calls same-event deduplication a poor fit after a canary found no duplicates.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks Could Help
Meyer’s proposal focuses on high-volume workflows where human review may be impractical and sending each item to a more capable model may add cost or delay. He describes Jev as a way to handle clear cases and route uncertain ones to another system or a person. The software using Jev’s answers determines that routing policy.
Language detection and comment triage are among the checks that could be applied across large collections. Meyer says his three live publishing checks processed close to 79,000 articles in one scan. The article does not provide independent verification or comparisons with other systems’ costs and accuracy.
Meyer recommends replaying 300 to 500 past decisions, reviewing disagreements and enabling a feature flag on a small canary before wider use. These steps are intended to evaluate performance before expanding automation.
Meyer’s Four-Part Fit Test
Meyer says a suitable task should have high volume, ask a narrow question without multi-step reasoning, have low-cost errors or a route for uncertain cases, and involve a heuristic that is visibly failing. He advises keeping a working keyword rule if it already performs well. In his framework, a use case is not proven merely because the question is easy or inexpensive to ask.
Before deployment, he recommends replaying historical decisions, comparing results overall and by confidence band, then reading a sample of disagreements to judge which system was right. He suggests wiring a use case into production only if its high-confidence results reach 95% accuracy in that evaluation. His stated benchmark for one 31-topic classification was 97% to 99% agreement with a frontier model at confidence of 0.8 or higher, versus 42% below 0.5. Those figures describe his measurement and agreement with that model, not necessarily accuracy against independently verified answers.
The article says its 24 examples span publishing, commerce, software, business operations and home use, but the supplied material details only the first nine examples. It identifies three live applications and gives publishing examples, while the rest of the map is not included in the available source text.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Evidence Still Needed for Seven Uses
Meyer says seven use cases need measurement first because the failure of the existing heuristic has not been demonstrated. The article excerpt does not supply results for those tests or describe the two poor-fit cases beyond same-event deduplication, which he says found no duplicates in a canary.
The reported benchmarks are Meyer’s own results. The provided material does not describe an independent evaluation, the full methodology behind the agreement figures, or how the results vary across different content and operating settings. It also does not give the other 15 use cases in enough detail to assess each one. These gaps leave open how broadly the reported performance and cost would transfer.
Measure Before Expanding Deployment
Meyer’s proposed next step for unproven applications is to replay 300 to 500 real past decisions, compare outcomes by confidence band and examine disagreements. He recommends enabling a separate feature flag that is off by default, testing on 5% to 10% of units, and expanding only after the high-confidence band reaches his 95% threshold.
The article does not announce a release date or deployment plan for the remaining examples. Whether the measure-first cases move into production depends on evidence that current methods fail and that Jev handles clear cases reliably.
Key Questions
What is Jev?
Meyer describes Jev as a tool that receives text or JSON and typed questions, then returns answers such as probabilities, category choices or scores for software to use in routing decisions.How many of the 24 uses are already running?
Meyer says three are live in his publishing operation. He rates 12 additional uses as strong fits, seven as requiring measurement first and two as poor fits.What does Meyer say Jev is best suited for?
His test calls for high volume, a narrow question, errors that are cheap or routed for review, and evidence that the current heuristic fails.Are the reported accuracy figures independently verified?
The available article material presents them as Meyer’s measurements. It does not describe an independent evaluation or establish that agreement with a frontier model is the same as accuracy against verified answers.What happens before a new use is deployed?
Meyer recommends replaying 300 to 500 past decisions, reviewing disagreements, and testing behind a feature flag on 5% to 10% of units before expanding.Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
