Diogo Almeida, a foundational figure in the development of conversational AI, experienced a profound disillusionment with the very technology he helped bring to prominence. As a key researcher at OpenAI, Almeida was instrumental in crafting ChatGPT and pioneering Reinforcement Learning from Human Feedback (RLHF), a technique that has become a cornerstone of modern AI training. Yet, despite the remarkable capabilities of these large language models (LLMs), he found himself increasingly frustrated by their perceived lack of true utility for automation.
"We have lightning in a bottle, and yet it is not useful," Almeida articulated in a candid interview, expressing a sentiment that has driven his subsequent work. "I’ve been battling that problem since then. It took me a while to come to the conclusion: The problem is we are optimizing for human language… We have been super good at human language for four years, but it’s not useful for automation because computers speak a different language."
This core conviction led Almeida to depart from OpenAI two years ago, embarking on a mission to rectify this perceived deficiency. He founded TypeSafe AI, a startup with the ambitious goal of building AI systems that communicate and operate on a fundamentally different principle. This week, TypeSafe AI unveiled its groundbreaking transformer-based model, Jev, a system that deliberately eschews the characteristics of traditional LLMs.
A Departure from Language: The Core of Jev’s Innovation
Unlike its language-centric counterparts, Jev does not generate human-readable text. Instead, its output consists of probabilities, which TypeSafe AI terms "calibrated decisions." This fundamental shift in output modality is the engine behind Jev’s distinct advantages. By sidestepping the complexities and ambiguities inherent in human language, Jev achieves remarkable efficiency and reliability.
The implications of this design choice are multifaceted and significant. Firstly, Jev operates at a fraction of the cost and speed of LLMs. The computational overhead associated with processing and generating natural language is substantial. By focusing on probabilistic outputs, Jev dramatically reduces the processing demands. Furthermore, the absence of text generation inherently eliminates the risk of "hallucination," a persistent challenge for LLMs where they generate factually incorrect or nonsensical information. In Jev’s architecture, users define the expected outputs in advance, ensuring that the model’s decisions are confined within a predetermined, reliable framework.
Another key economic advantage lies in its tokenomics. While LLMs typically meter both input and output tokens, Jev’s output tokens are essentially free, with input tokens being metered by the billion rather than the more restrictive million. This pricing structure makes Jev an exceptionally cost-effective solution for tasks requiring high volumes of decision-making or analysis.
Early Adopters and the Promise of Automation
The response from the developer community to Jev has been overwhelmingly positive, with early indicators suggesting a significant demand for its capabilities. The company experienced a temporary outage of its API service due to an unanticipated surge in user traffic, a testament to the immediate interest Jev has garnered. The primary application area where Jev is demonstrating its most profound impact is in software automation. Developers are recognizing Jev as a more robust and economical method for integrating intelligent decision-making into their applications.
Concrete examples are already emerging. Pranit Sharma, a software engineer at Vercel, a company specializing in agentic infrastructure, shared his team’s experience integrating Jev. Vercel had previously employed OpenAI’s ChatGPT Luna 5.6 for a crucial task: reviewing commands for safety through a classification process. Upon replacing ChatGPT Luna with Jev, Vercel observed a dramatic improvement, achieving results five to eighteen times faster and with demonstrably greater accuracy. This substantial performance leap underscores Jev’s efficacy in specialized, high-throughput automation tasks.
Further validation comes from Nikhil Mudholkar, CTO of Bryo AI. In a comparative test designed to classify business emails, Mudholkar pitted Jev against Google’s Gemini. While Gemini exhibited slightly higher accuracy in this specific instance, the cost differential was stark: Gemini proved to be ten to twenty times more expensive. Crucially, Mudholkar highlighted Jev’s confidence scores as a game-changer. "It is the only one that hands back a real probability which makes it ideal for automating workflows!!" he enthused, emphasizing the value of quantifiable certainty in automated processes.
Beyond Replacement: Jev as an Augmentative Tool
Jev’s utility extends beyond merely replacing LLMs in specific use cases. It also offers significant potential as an augmentative layer, acting as a critical safeguard against the misbehavior of other AI systems, particularly LLM-based agents. The proliferation of AI agents, each potentially requiring oversight, can quickly escalate operational costs. However, Almeida posits that deploying Jev for this monitoring role presents a compelling economic and functional solution. He envisions Jev being used to scrutinize the operational traces of LLM agents, thereby preventing unauthorized actions or "jailbreaks" where models are manipulated to bypass their safety constraints.

Armin Ronacher, CTO of Earendil, a company that develops the open-source model harness Pi, elaborated on this point. "At the end of the day, it delegates the hallucination problem a little bit to the user," he explained. "The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it." This perspective highlights Jev’s role in empowering users with granular control over decision-making based on calibrated confidence levels.
Another promising application identified by Ronacher is model routing. In complex AI systems, determining the most appropriate model for a given task is crucial for efficiency. Utilizing an LLM for this predictive routing function would be prohibitively expensive. Jev’s low cost and high speed make it an ideal candidate for real-time, intelligent workload distribution, ensuring that tasks are handled by the most suitable and cost-effective AI component.
The Jevons Paradox and the Democratization of Intelligence
The naming of Jev is not coincidental; it is a deliberate nod to William Stanley Jevons, a 19th-century economist whose eponymous paradox observed that increased efficiency in the use of a resource can, counterintuitively, lead to increased consumption of that resource. Almeida hopes that Jev, by dramatically reducing the cost of intelligence, will similarly lead to its widespread and pervasive deployment.
"We think that there’s just going to be smart software all over the place in a way that’s emergent and distributed… much more like the early internet than you know like the mega apps that people are trying to build right now," Almeida stated, painting a vision of a future where intelligent capabilities are seamlessly integrated into the fabric of digital systems, akin to the decentralized growth of the early internet.
While Almeida remains circumspect about the specific architectural details of Jev, outside observers speculate that it may be built upon an open-weight LLM foundation. TypeSafe AI categorizes Jev as a "System One model," emphasizing its intuitive decision-making capabilities rather than complex reasoning. The company’s blog posts highlight a focus on "the right task," suggesting a specialized approach to AI development.
A significant aspect of Jev’s development is its training methodology. Almeida revealed that Jev is trained exclusively on synthetic data, utilizing a novel technique he terms "reinforcement learning from calibrated decisions." This approach diverges sharply from traditional reliance on vast quantities of real-world, human-generated data.
"We made an early bet that we will be making all of our data, and that has been one of the best bets I’ve ever made in my life – better than our launch, in my opinion, better than RLHF," Almeida shared. He further elaborated that a substantial portion of TypeSafe AI’s operations is dedicated to a lab focused on statistically well-understood synthetic data, a domain that now brings him significant professional satisfaction.
The Dawn of a New AI Paradigm?
Currently, Jev stands as a singular offering in its class. However, as its utility becomes increasingly apparent, industry experts like Ronacher anticipate the emergence of competitors. "We should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don’t have to be creative yet," he commented, suggesting that the current economic landscape of LLMs may have masked the need for alternative approaches.
TypeSafe AI is committed to expanding Jev’s capabilities, with plans to develop versions of the model that can operate across new modalities. When asked about the company’s positioning, Almeida expressed a desire to avoid the hype often associated with "frontier labs." "The main product of frontier labs is fear or hype. I would like our main product to be intelligence… [but we are] not a lab in the sense of, you know, like bet on infinite wealth, or a religion, or building God in a data center, or whatever is the thing of today." This statement underscores a grounded and practical vision for AI development, focused on tangible utility and measurable impact.
The evolution of AI is clearly at a crossroads, with foundational researchers like Diogo Almeida challenging the prevailing paradigms. Jev represents a bold step in a new direction, prioritizing efficiency, cost-effectiveness, and reliability for automation by stepping away from the seductive but often impractical allure of human language. Its success could herald a significant shift in how intelligent systems are developed and deployed, moving towards a more distributed, task-specific, and ultimately, more useful form of artificial intelligence.







