Latent Space: The AI Engineer Podcast Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week 1 1× 0:00 Current time: 0:00 / Total time: -39:12 -39:12 Audio playback is not supported on your browser. Please upgrade. Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week Our DevDay coverage - the first pod on the DevDay lineup - dives in with the leaders of OpenAI’s CUA team and API platform. Sep 30, 2026 1 Share Transcript Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: Dwarkesh Patel @dwarkesh_sp Here's a question I find confusing and interesting and which actually tells us a lot about the nature of current AI progress: Why has progress on computer use been so slow? Computer use is so clearly verifiable. I think the answer is that it is not enough for a domain to be… Dwarkesh Patel @dwarkesh_sp What does the next training paradigm look like? 0:00:00 – The big research bet the labs are making 0:02:12 – Grindability is just as important as verifiability 0:06:10 – Will RLVR alone generalize? 0:08:41 – Getting the learning back to the weights 0:15:22 – Dreaming 0:17:23 – 12:54 AM · Jun 27, 2026 · 604K Views 143 Replies · 70 Reposts · 1.02K Likes We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch , there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein , cofounder of Sky and now leading all the amazing CUA progress that casuals might miss: Ari Weinstein @AriX We've gotten a lot of great questions on privacy implications of Computer History. Here's a few things we've done to build a super powerful feature while keeping things private. First off, you can review all of your Computer History in the timeline view: Ari Weinstein @AriX Today we're releasing Computer History in ChatGPT. It lets ChatGPT learn from everything you do on your computer, so it can better understand how you work, finish tasks that you're in the middle of, and suggest skills and automations based on how you use your computer. 8:39 PM · Aug 14, 2026 · 107K Views 39 Replies · 29 Reposts · 512 Likes Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software. OpenAI clones Jev In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API . Given that we were the first Jev podcast , we particularly focus on the unusually fast sprint on the Decisions API: OpenAI Developers @OpenAIDevs Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna. Define questions and possible answers to classify content, route requests, or choose an agent’s next action. Available in limited preview. 6:34 PM · Sep 29, 2026 · 684K Views 222 Replies · 401 Reposts · 5.71K Likes And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns. We discuss: Why OpenAI thinks Computer Use has changed dramatically in just the last few months Dots and what changes when every agent gets its own Linux computer Why Computer Use can now complete some tasks faster than the average human The path from human-level to “literally superhuman” computer use Why modern agents are much better at debugging and recovering from failure How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together App Shots and why they give models much richer context than ordinary screenshots Why Computer Use can close the loop between writing software and testing it Trust, permissions, and safety when agents can make payments and operate websites Async function calling and why models no longer need to stop reasoning while tools run Mid-turn steering, WebSockets, and the architecture behind more responsive agents UltraFast inference and how OpenAI is pushing frontier models toward much lower latency The rapid internal story behind the Decisions API Why Decisions API is more than structured outputs at low latency GPT Live, fast tool calling, and real-time computer control How OpenAI is already using Decisions API for support classification and internal workflows Longer prompt caching, cache pre-warming , and cache-aware applications Server-side compaction vs manual compaction for long-running agent threads What should live inside an Agents API versus a developer’s own harness OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs Ari Weinstein Product & Engineering, Computer Use at OpenAI X: https://x.com/AriX LinkedIn: https://www.linkedin.com/in/weinsteinari/ Nikunj Handa Product, API at OpenAI X: https://x.com/nikunjhanda LinkedIn: https://www.linkedin.com/in/nikunjhanda/ Timestamps 00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API 00:02:52 Dots and Personal Cloud Computers 00:04:59 Why Computer Use Is “180 Degrees Different” 00:06:04 From Sky to Self-Debugging Computer Use Agents 00:09:24 How Computer Use Sees and Operates Software 00:12:09 From Faster Than Humans to Superhuman Computer Use 00:16:03 Agents API: Trust, Permissions, and Safety 00:17:31 Computer Use for Coding, Testing, and QA 00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference 00:23:21 The Rapid Story Behind Decisions API 00:25:32 What Decisions API Is and How It Works 00:30:24 What OpenAI Is Building With the New APIs 00:32:23 Prompt Caching, Pre-Warming, and API Performance 00:35:20 Context Compaction for Long-Running Agents 00:37:13 Memory, Higher-Level APIs, and the AI Cloud Transcript Introduction: OpenAI DevDay and the New Agent Stack Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast Swyx [00:00:08]: We’re the first podcast after your livestream. Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today? Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it. Swyx [00:01:02]: And the API. Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote. Swyx [00:01:35]: And not to mention the Decisions API. Ari Weinstein [00:01:37]: Decisions API. Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models? Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate? Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast. Dots and Delegating Work to a Cloud Computer S
Latent Space: The AI Engineer Podcast Why Dwarkesh is Wrong about Comp
Latent Space: The AI Engineer Podcast Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week 1 1× 0:00 Current time: 0:00 / Total time: -39:12 -39:12 Audio playback is not supported on your browser. Please upgrade. Why Dwark
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。