Baseten 工程博客(网页) · 2026/10/3

Model performance Agentic inference optimization: 50-90% faster engine

Model performance Agentic inference optimization: 50-90% faster engines Claude Code with Fable 5 built a Qwen 3.6 inference engine that beat vLLM by up to 90% on decode speed and 2.3x on TTFT, using MetaInfer skills on one B200. Authors Shawn Rushefsky Last up

Model performance Agentic inference optimization: 50-90% faster engines Claude Code with Fable 5 built a Qwen 3.6 inference engine that beat vLLM by up to 90% on decode speed and 2.3x on TTFT, using MetaInfer skills on one B200. Authors Shawn Rushefsky Last updated October 2, 2026 Share Baseten is the kind of place where we have Slack channels like #ai-papers-discuss and #model-performance-reading-group , and it was in one of these channels that a paper caught my attention recently: MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox . In it, the authors provide a skills-only framework for building custom inference engines from scratch, claiming their results beat SoTA open solutions like vLLM on common inference performance metrics like time per output token (TPOT). “Skills-only” here means no model post-training, and no specialized harnesses. I’ll admit, I was skeptical, but they provided a repo with the skill toolkit, and we’re still living that unlimited-token good life here, so I thought I’d give it a try on a popular model just to see how it did. With Qwen-3.6-35B-A3B in NVFP4 precision on a single B200, our LLM-generated inference engine outperformed vLLM 0.25.1 (latest at the time of the experiment) by up to 90% on single-stream decode speed, and with 2.33x faster TTFT, opening up new ultra-low-latency use cases for the model that had not been practical before. The rise of AI-assisted inference optimization There are a few large-scale trends that are converging right now, creating fertile ground for discoveries and techniques like MetaInfer. ✕ Four converging trends: smarter models, agent coordination, measurable optimization, and inference demand. Models are getting smarter and better at long-horizon tasks, with the release of Anthropic’s Mythos 5 and Fable 5 models marking a meaningful step forward on the kind of tasks that can be reliably delegated to AI systems. The last such moment I remember was the release of Opus 4.5 in late 2025, at which point many engineers, myself included, moved from using AI coding tools as a spicy autocomplete to driving coding harnesses like Claude Code as our primary work interface. At the same time, agent coordination workflows and primitives are getting packaged directly into popular harnesses, making complex multi-agent workflows more accessible and easier to operate. One important such feature is “Goals”, provided by both Claude Code and Codex, which keep models working and on-task toward measurable goals over long periods of time. I first encountered this experiment-measure-iterate loop in Karpathy’s Autoresearch project about 6 months ago, but it's now just part of normal coding harnesses. ✕ AI inference demand is compounding. Source: Google I/O 2026. As the capabilities of AI systems keep increasing, so does the market demand for inference services, creating enormous pressure to continuously improve both cost and performance. With all of these factors combined, it starts making a lot of sense to point LLMs at well-defined optimization problems and just let them grind away at solutions. Inference performance turns out to be an ideal domain for this kind of automated optimization. The problem space is well-defined in terms of quantifiable metrics like TTFT, TPOT, throughput-per-chip, and required memory. Even correctness can be numerically verified against full-precision reference implementations of the model, to ensure that performance gains do not come at the expense of accuracy. This objective verifiability is critical to the success of autonomous optimization. We also know that there is a huge gap between current SoTA performance and speed-of-light theoretical maximum performance. Additionally, inference is enormously complex, and the search space for optimizations is very large, ranging from low-level CUDA kernel authoring all the way to serving-layer optimizations like dynamic batching. Constraining the solution space LLMs may be very good at solving well-defined optimization problems, but now the challenge shifts to how to frame your problem as a well-defined optimization problem. You need to be able to define your constraints well, and choose your optimization targets carefully. The agent will absolutely try to game the system to “pass” the goal, so it’s up to you as the engineer to design the game in such a way that the agent’s success actually aligns with what you need. As an example, if you don’t set accuracy-preserving constraints, your “model” will converge on: while True : yield "a" Such a simple model will have unbelievably good TTFT and TPOT scores, but is no longer useful for anything. Existing open inference engines like vLLM, SGLang and TensorRT-LLM do an overall great job at serving a wide range of models on a large variety of accelerators at commercially acceptable performance levels. However, the benefits they offer in generality and ease of use come with a tradeoff: they are not hyper-optimized for any specific (model + accelerator + workload) tuple. The premise behind MetaInfer is that for any specific deployment, a highly specialized (i.e., constrained) engine can outperform these generalized engines. Adapting MetaInfer for production applications In the original MetaInfer paper, they challenge off-the-shelf agents to build inference engines from a knowledge base of contracts and constraints, with no direct access to source code from open engines. It’s extremely likely that such source code was present during model training, but agents did not have direct file-level access. Agents expand the knowledge base as they work, proactively identifying gaps and filling them through structured experimentation and verification. While the experimental purity taken on by the authors is admirable and interesting, we’re very outcome-oriented at Baseten. I did allow my agents to access SoTA open solutions for reference when needed, including finding pre-optimized kernels from vLLM and TensorRT-LLM where available, and only writing their own kernels when needed. Additionally, while the original paper focused on just the inference engine component, I broadened the scope to the full serving stack, with final outcomes measured against deployed multi-replica services behind a load balancer, using AIPerf to generate load and measure performance. In both the original paper and my extension, the result is two deliverables: the inference engine itself, and reusable extensions to the knowledge base that can be used in subsequent runs. The first experiment: Qwen-3.6-35B-A3B inference Setup In kicking off the experiment, I gave Claude Code with Fable 5 access to the paper and the MetaInfer Repository , as well as access to an NVIDIA B200 workstation via SSH. I told it the target hardware and provided a link to the Hugging Face repo for both the NVFP4 quantized weights I wanted to use for Qwen-3.6-35B-A3B, as well as the original full-precision weights to be used as an accuracy oracle. The agent was also able to deploy engine candidates to Baseten, and to run AIPerf profiles against the deployed endpoint. We have internal agent skills and MCPs to facilitate all of the required Baseten interactions, and the agent did have access to those as well. I set a /goal telling it I wanted to beat vLLM by 20% on all performance metrics without accuracy loss vs. the NVFP4 baseline. I was surprised and impressed both by the final results and the relatively low amount of human intervention required. Claude worked mostly autonomously for roughly 1 week on this project, with me providing occasional steering to keep it from over-fixating on one particular traffic shape or another, and to keep it focused on production inference performance, and not just isolated engine benchmarks. The MetaInfer framework uses a series of immutable gates and checks that enable productive long-horizon work like this. Claude would periodically pause to ask me to approve any changes that came with accuracy differences. I ultimately allowed it to have subtle numerically different outputs from the reference NVFP4 implementation, as long as overall accuracy vs. the BF16 baseline was at least as good. Results I did very little to manually manage agent coordination or context management other than requesting that it use Fable 5 for all subagents after some disappointing early results from smaller models. It underwent automatic context compaction many times, and consumed roughly 1.7 billion tokens (overwhelmingly cached input tokens) and ~200 B200 hours. It actually reached parity with vLLM relatively quickly, within the first few days, but I had it keep grinding away for the sake of science. By the end, the generated engine, called VibeQwen , outperformed vLLM across all tested traffic patterns, and by very large margins in some cases. In head-to-head comparisons on a single B200, VibeQwen achieved 1,792 tokens per second (TPS) on single-stream, speculator-friendly text (repetitive and structur

FDE 判断

这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。