🔍 Read the full analysis: Jev For AI Decisions: 24 Methods, Workflows, And Use Cases on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 potential uses for Jev, a system that returns typed answers to narrow questions so software can act on them. Meyer says three publishing workflows are live, 12 other uses meet his fit test, seven need measurement and two are poor fits; those figures and performance results are his own reports.
Meyer describes Jev as a system that receives text or JSON state alongside typed questions and returns answers that code can branch on, rather than prose to interpret. The answer types include a yes-or-no probability, a choice with probabilities and confidence, or a score on ordered levels. He reports that one call takes about **0.3 to 0.9 seconds** and costs about **$0.04 per million input tokens**.
In his publishing operation, Meyer says a relevance gate assessed about **10,000 story-site pairings in three days**, with 22% judged clearly on-topic. A language check scanned 78,889 articles for $2.01, found 1,576 non-English items and fixed 1,553. A classifier fallback agreed with a frontier LLM 89% of the time overall, according to Meyer; in a 31-topic measurement, he says agreement was 97% to 99% at confidence of 0.8 or higher and 42% below 0.5.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test Before Adding Automation
Meyer’s proposed test for automation requires high volume, a narrow decision, errors that are cheap or routed to a stronger reviewer, and evidence that an existing heuristic visibly fails. He recommends retaining a working keyword rule unless measurement shows a gap.Meyer recommends replaying **300 to 500 past decisions**, comparing results overall and by confidence band, then reviewing 20 disagreements to determine which system was right. He says to use a workflow only where the high-confidence band reaches 95%, with a separate flag, disabled by default, and a canary on 5% to 10% of units.
In the relevance gate, Meyer says low-fit cases are dropped only when Jev is confident; uncertain cases continue to publish as before. The article reports these figures as Meyer’s results and does not provide an independent evaluation.
How Jev Fits Publishing Checks
Meyer presents the three running workflows as examples of checks made practical by low-cost, repeated decisions. The **relevance gate** combines two yes-or-no questions and a fit score to compare a story with a site profile. The **language check** asks whether an article title and opening paragraph are in English; Meyer says items below a 0.1 threshold are rewritten in place at the same URL. The **classifier fallback** chooses among 31 topic descriptions when the primary LLM errs, with a keyword rule as a last resort.His fit labels distinguish deployed uses from potential ones. “Live” means running in his fleet; “strong fit” means all four conditions are met; “measure first” means the failure of the existing heuristic has not been proven; and “poor fit” means at least one condition fails. In publishing, he lists **disclosure detection** and **comment moderation** as strong fits. Product matching in roundups, headline quality and thin-source detection need measurement first. Same-event deduplication is a poor fit in his canary because he found zero duplicates.
The examples show how the test applies to existing rules. Meyer says a 300-character rule cannot distinguish a fact-dense wire item from a teaser, while a regex may miss paraphrased disclosures about free products or affiliate links. For disclosure checks, his suggested rule sends misses to human review and bars automatic publication. For moderation, he proposes auto-approving acceptable comments or hiding spam only above a 0.9 confidence threshold, with other cases queued.
““Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.””
— Thorsten Meyer, in the Sept. 29 article
Evidence Still Needed Across Uses
The reported performance figures come from Meyer’s own measurements; the source does not describe an independent audit, detailed evaluation method or comparison dataset for the production workflows. The reported agreement with a frontier LLM is a comparison with that model, not a measure of verified correctness against ground truth. The article also does not establish how results vary across different publishers, content types or operating conditions.Seven of the mapped uses still need evidence that the existing heuristic fails, by Meyer’s own classification. The source says his deduplication canary found zero duplicates, but gives no sample size or time window. It also ends during the commerce and customer operations section, leaving the remaining use cases and their specific rules unavailable in the supplied material. The stated total of 24 can be reported, but the full inventory cannot be reconstructed from this excerpt.
Measure Before Wider Rollout
Meyer recommends that prospective users replay **300 to 500 historical decisions** and examine 20 disagreements before enabling a workflow. If the high-confidence band meets his 95% threshold, he recommends putting the integration behind its own flag, beginning with a 5% to 10% canary and expanding from there. These are recommendations in his article; the source does not announce a broader product launch or a scheduled follow-up milestone.For the seven “measure first” cases, the next evidence would be a measured error rate for the current rule or matcher. Until that is available, Meyer’s framework classifies the cases as promising but unproven. The supplied article excerpt does not say when he will publish the complete list or provide further results.
Key Questions
What did Thorsten Meyer publish?
He published a guide mapping **24 potential Jev use cases** across publishing, commerce, software, business operations and the home. The provided source excerpt details publishing cases and only begins its commerce section.
How many use cases does Meyer say are ready or running?
He says **three are live** in his publishing operation and 12 additional cases are strong fits. Seven need measurement, while two are poor fits. These classifications are Meyer’s assessment.
What does Jev return?
According to Meyer, Jev returns typed answers such as yes-or-no probabilities, choices with confidence, or scores on ordered levels. Software can use those outputs to branch without parsing a prose response.
What evidence does Meyer recommend before deployment?
He recommends replaying **300 to 500 real past decisions**, checking results by confidence band and reviewing 20 disagreements. His proposed rollout threshold is 95% in the high-confidence band, followed by a limited canary.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
