Dev & Tools
Frameworks, model releases, engineering deep-dives and practitioner notes.
Jev introduces a new shape of LLM - System One, aka Decision Models
Last week TypeSafe AI unveiled Jev , their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores. TypeSafe describe Jev like this: Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. It's also very fast, and really cheap . Regular LLMs are priced in terms of input and output tokens, with output generally charged at significantly higher rates. Jev charges only for input - output is free - and the input price of their first model is $0.042 per million tokens - cheaper even than OpenAI's GPT-5 Nano ($0.05/million). Jev lets you ask questions about text or semi-structured data. You compose a "state" object containing a string, array of strings, or set of name-value pairs - this might describe an article, or a customer, or any other kind of record. You then send that to their API with one or more questions, and get a reply back for each. You can ask three kinds of questions: Yes/No questions, which Jev calls "Noul" questions - their CEO confirmed on Hacker News that this is short for Bernoulli, from the Bernoulli distribution . You pose a statement and get back a floating point number between 0 and 1 for how confident the model is that the statement is true. Choice questions, where the model picks one from a set of provided options - actually a confidence score plus a probability distribution across all of the options. Score questions, where you provide sequence of numeric levels with descriptions and it provides a floating point score somewhere along that range. The Jev API can accept a single document ("state") and as many questions as you can cram into the context window. Questions are evaluated in parallel, so sending many questions should take a similar time to sending just one. The Jev 1.13 jaggedness documentation offers useful guidance as to Jev's strengths and weaknesses. It's currently not great with numbers, dates, or "adversarial content". I think the decision model framing is useful for understanding where to use Jev. It's great for anything that can be expressed as a classification task - think spam detection, suggesting labels, prioritization and
Roundtables: The Deadly Failures of The Virtual Border Wall
The US has spent billions building a “virtual wall” of surveillance towers along its southern border over the past 25 years, promising they will help detect and apprehend border crossers and save lives. But a groundbreaking investigation by MIT Technology Review has documented over a thousand people who moved through areas watched by these towers without being reached or apprehended, and who ultimately died there. Some were even under the watch of newly installed AI-powered towers designed to spot people automatically. Our findings reveal a humanitarian crisis more visible than previously known, and repeated failures of the virtual wall’s basic security promise. Join MIT Technology Review editors and reporters for a conversation examining the failures of border surveillance technology and uncovering the stories of the people who die in the borderlands. REGISTER NOW Speakers : Mat Honan , editor in chief, James O’Donnell , senior AI reporter, and Eileen Guo , senior features and investigations reporter Related Stories The US spent billions on border surveillance. Why can’t it catch people before they die? How we made the first comprehensive map of deaths along the US border’s “virtual wall” 4 ways to address the failures we found along the US border’s “virtual wall”
Dev & ToolsThe US spent billions on border surveillance. Why can’t it catch people before they die?
When José Morales Bernal crossed the border into the United States on April 8, 2024, the day before his 32nd birthday, it should have triggered a chain of technological alerts and human responses. As he walked through the desert in southern New Mexico that morning, he was within range of three surveillance towers. Newly installed by US Customs and Border Protection (CBP), they were built by the defense tech company Anduril and equipped with cameras and AI to automatically detect and track people. They transmit live video to nearby control rooms and can send alerts to the government-issued smartphones held by agents in the area, prompting the closest available to respond. This story is part of Dying on Camera , a collaboration between MIT Technology Review and Times of San Diego . Journalists in both newsrooms spent the past year examining the failures of border surveillance technology and uncovering the stories of the people who die in the borderlands . These AI-enabled towers are meant to give greater visibility across the 1,951-mile southern border, freeing up border agents from having to spend hours staring into video monitors. They were installed in this particular place to spot border crossers before they reached the nearby town of Sunland Park. If the system worked as intended, Morales should have been apprehended. If he needed medical help, agents were trained to provide it. The view from where José Morales Bernal’s body was found toward a Customs and Border Protection surveillance tower made by Anduril, along the southern US border in Sunland Park, New Mexico. CENGIZ YAR FOR MITTR That didn’t happen. Despite the nearby surveillance towers, it was employees of the local landfill, rather than Border Patrol, who first spotted Morales that morning. At 1 p.m. the workers saw him again, now lying in the sand. At 4 p.m. the landfill workers saw that he had not moved and called Border Patrol. Agents arrived 45 minutes later. He was dead. When an agent then called 911 to report the body, he said it was “probably one of the migrants crossing through there,” seemingly unaware that Morales had been moving near the agency’s surveillance systems earlier that day. The agent can be heard sounding unsure of Morales’ location despite the proximity of an AI-powered surveillance camera. The call has been edited for length. Morales had died just 360 feet from the closest surveillance tower. Two more towers stood watch to the east and the west. An autopsy later co
Dev & ToolsFive AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda
Would you trust an AI agent with access to your company’s systems? Put your employees in an autonomous vehicle? Deploy a robot that has to make sense of an unpredictable physical world?
Dev & Tools4 ways to address the failures we found along the US border’s “virtual wall”
MIT Technology Review today published our investigation into how many people have died near the “virtual wall” of surveillance towers that the US government has installed along the US-Mexico border. We found cases of people who walked undetected through areas surveilled by advanced, AI-enabled towers and later died nearby, where their bodies remained unnoticed for weeks or months. These cases reflect systemic flaws in the virtual wall: broken towers, algorithms that failed to detect people, and agents who simply did not respond to alerts. This story is part of Dying on Camera , a collaboration between MIT Technology Review and Times of San Diego . Journalists in both newsrooms spent the past year examining the failures of border surveillance technology and uncovering the stories of the people who die in the borderlands . Each of these deaths should have raised alarms inside US Customs and Border Protection (CBP) about the flaws in the agency’s surveillance network. This is particularly true for cases near the latest towers built by the defense company Anduril, which continues to win major contracts and is poised to benefit from the windfall of federal money slated for border security. The US is set to spend $1 billion to triple the size of the virtual wall by 2034. Here are four ideas for what CBP should do next to address these failures. 1. Conduct a comprehensive audit of deaths near the virtual wall In this June 16, 2021, file photo, a family from Venezuela moves to a Border Patrol transport vehicle after they and other migrants crossed the U.S.-Mexico border and turned themselves in to agents in Del Rio, Texas. ERIC GAY/AP PHOTO “If I were the federal government, I would do an audit, because this feels like it’s not working.” That was New Mexico state representative Sarah Silva’s reaction when MIT Technology Review showed her that in 2024 alone, more than 18 people had died near surveillance towers in just one small stretch of desert near her district. An audit is necessary because, despite poring over thousands of pages of police reports and other documents about surveillance towers, we have only limited knowledge of how many towers are out there and where they’re located. CBP does not publish a comprehensive map of its virtual wall. We instead relied on mapping work led by the nonprofit Electronic Frontier Foundation and then used satellite imagery to verify where specific towers were and approximately when they were installed. We analyzed lots of to
Dev & ToolsGemini Hacked Three Companies in First Known Breakout by Google’s AI
Gemini Hacked Three Companies in First Known Breakout by Google’s AI Gemini finally caught up on Felony Bench ! The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta. In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems. In each case, the model ended the intrusion after determining it had accessed a real company’s systems, Google said. Gemini is apparently less determined than other models, and decided not to keep going. Google knew about these in July, but chose not to disclose them until the WSJ reached out, presumably based on a tip. Google said it didn’t consider the hacks to warrant public disclosure—because its model didn’t cause harm to the companies and ended each intrusion immediately upon determining it had hacked a real company rather than a simulated one. Tags: security , ai , generative-ai , llms , gemini , accidental-cyberattacks
Dev & Tools🤖 AI Agents Weekly: Jev, Salesforce Koa, Claude Code Projects, Anthropic R&D Metrics, Periodic Neon, Gemini 3.8 Live, and More
In today’s issue: TypeSafe launches Jev Salesforce releases Koa Claude Code ships Projects Anthropic publishes AI R&D metrics Periodic Labs releases Neon Gemini 3.8 Live launches Devin adds codebase-wide Code Scans HarnessTax measures the harness cost Claude Code reads AGENTS.md Cowork merges into Claude And all the top AI dev news, papers, and tools. Top Stories Jev and System One Models TypeSafe AI came out of stealth with Jev, a model built for the small decisions software makes millions of times a day, not for chat. Founder Diogo Almeida worked on the instruction-following methods behind ChatGPT at OpenAI. Typed decisions, not text: Jev takes your app state plus a typed question and returns a decision with a calibrated probability attached. You never write a JSON prompt, add a parsing layer, or validate the output. Three question shapes: A boolean question returns a probability, a choice question picks one option from a set you define, and a score question returns a number on your scale. You can run several questions about the same input in one call. Speed and price: Answers come back in 70 to 500 ms end-to-end at $0.042 per million input tokens, with output tokens free. TypeSafe reports up to 190x faster and 440x cheaper than frontier LLMs on its published workflows. Error rates: Jev records 0% structured output errors and 0% tool call errors on TypeSafe’s suite, against 5.73% and 0.67% for Opus 5. Training method: The model uses a new architecture, a parallel sampler, and a training method TypeSafe calls Reinforcement Learning for Calibrated Decisions. Blog | OpenRouter Salesforce Koa Salesforce released Koa, a 120B enterprise model post-trained from NVIDIA’s open-weight Nemotron-3-Super-120B with GRPO, aimed at multi-turn tool use in CRM workflows. Specification-driven RL: A simulation-to-reward pipeline expands workflow specifications into persona-conditioned multi-turn tasks, with task-resolution rewards grounded in successful tool use for data-dependent requests. Enterprise specifications are written in Agent Script, Salesforce’s declarative language for building Agentforce agents, and public tool-use specifications are synthesized directly. Tool-use results: 66.63% on BFCL against 64.73% for the Nemotron base and 53.96% for GPT-4.1. On Tau2Bench, it reaches a task-weighted 69.41 against 68.64 for the base and 54.48 for GPT-4.1, still behind Opus 4.8 at 74.00 and GPT-5.5 at 83.99. CRM Bench: 0.86 overall against 0.84 for t
Dev & Toolsllm-typesafe 0.1a0
Release: llm-typesafe 0.1a0 I built this new plugin for LLM to add support for TypeSafe AI's new Jev model . Install it like this: llm install llm-typesafe Then set an API key ( get one here , the waitlist seems to move pretty fast): llm keys set typesafe # Paste key And now you can ask yes/no "noul" questions like this: llm -m jev 'Please refund my last payment.' \ -s 'Does this message explicitly request a refund?' Output: {"type": "noul", "noul": 0.99} Or choice questions like this: cat message.txt | llm -m jev \ -s ' Which team should handle this message? If billing and technical issues both occur, choose billing. ' \ -o answer_type choice \ -o criteria ' { "billing":"Charges, invoices, payments, or refunds", "technical":"Problems installing or using the product", "other":"Neither category fits" } ' Or scoring questions like this: cat report.txt | llm -m jev \ -s ' How reproducible is the problem described in this report? ' \ -o answer_type score \ -o criteria ' [ "No reproduction instructions", "Some instructions, but important steps are missing", "Complete steps with expected and actual results" ] ' See the README for more details. Tags: projects , llm , jev