Latent Space · 2026/10/7

AINews: Weekday Roundups [AINews] Quasi-Riemann-Hypothesis: OpenAI pub

AINews: Weekday Roundups [AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics Our head hurts. Oct 07, 2026 ∙ Paid 11 Share Tickets for AIE N

AINews: Weekday Roundups [AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics Our head hurts. Oct 07, 2026 ∙ Paid 11 Share Tickets for AIE NYC are selling out soon! See you next week! see past AINews issues for subscriber discounts. Pour one out for Mistral, who shipped a decent Large 4 “Le Chonk” model on the new 3800 GB300 cluster funded by their recent Series D . But they were overshadowed by more mathematics results from OpenAI’s internal Navier-Stokes math model - published as a blogpost , repo , and tweet . The best compliment comes from their Navier-Stokes competitor from Anthropic, who despite his personal issues with Anthropic, does not mince words: “It’s obviously the most significant moment in mathematical history.” levent @ __alpoge__ Big, big, big, big props for quasiriemann and no Siegel zeroes (with many other beauties in there), I’m kicking myself for talking all over about it being within reach but not actually having pushed. There are some sad stories related to their users getting scooped / conflicts… 11:39 PM · Oct 6, 2026 · 99.4K Views 29 Replies · 100 Reposts · 1.63K Likes This bears some qualification , but most experts seem to agree that it solves many of the top 500 open problems in math . In particular, Result 003, the Quasi-Riemann Hypothesis , is somewhere between a Fields Medal result and “the biggest result in number theory in 200 years ”. The most astonishing is the how - while Navier-Stokes was done in 88 hours and 10,000 agents , these solutions were 3 hours of ChatGPT Pro on average . AI News for 10/5/2026-10/6/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap OpenAI Releases 722 Math Manuscripts From an Unreleased Internal Model The release : OpenAI published a broad set of mathematical results from an internal frontier model in a public GitHub repo . It says it consulted the Institute for Advanced Study’s independent Advisory Group on Mathematics and AI on how to release them. Scale and compute : The collection reportedly holds 722 manuscripts grouped into 372 families of related results. They came from an evaluation of about 4,000 research problems and used an average of roughly three hours of ChatGPT Pro thinking compute per result ( summary , Rundown ). Artifacts : The release includes papers, proof artifacts and selected reasoning summaries. The model itself remains unreleased. Framing : Sam Altman called it “a new era of discovery” . Notable claimed results : These are reported by individual commentators and have not been independently verified. Integer multiplication : One contributor highlighted a result for integer multiplication faster than n log n . Elastic inverse problem : Another singled out a uniqueness result for the elastic inverse problem , which the paper says had been open in 3D since 1994. Millennium-adjacent work : Commenters point to partial progress on Riemann, Hodge and BSD . Mathematician reaction : Levent Alpöge praised the quasi-Riemann and no-Siegel-zeros results and called it “the most significant moment in mathematical history” . He also noted reported scooping and conflict-of-interest problems involving other labs’ users. Composition of results : An analysis estimates about 20% of the results are disproofs or counterexamples . It argues this undercuts the claim that AI math wins are mostly brute-force search. Skepticism and open questions : Errors expected : Will Depue expects that some results should not survive scrutiny . He built citedbyagi.com to track which human papers the release cites. Compute framing : Teortaxes notes that three hours of compute “is not much” . Generalization : François Chollet asks whether gains in RLVR-friendly math and code generalize, or whether non-verifiable domains stay bottlenecked on human data . Mistral Large 4 (”Le Chonk”): Launch, Pricing and Contested Evals Mistral Large 4 preview : The model has 1T total parameters and 49B active, is natively multimodal and is available via API now ( announcement ). Open weights are promised for end of October. Training status : The RL run is “still in flight and shows no sign of saturation” . Compute : The model was pre- and post-trained on ~3,800 Grace Blackwells in Europe . A larger model is training now . Pricing : $1.36/$4.18 per million input/output tokens, with $0.14 for cached input and 50% off for the first two weeks ( Artificial Analysis ). Context : Vals and Artificial Analysis list a 512K context window. OpenRouter lists 1M context with up to 256K output . Mistral’s own claims : Human evals : Mistral says it beats GLM 5.3 on STEM, CAD and finance in human evals and is on par in agentic coding. Coding benchmarks : It reports outperforming GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4 ( Rozière ). Blind review : In a blind Surge coding review it finished #2, behind only Opus 5 . Independent measurements : Artificial Analysis : It scores 38 on the Intelligence Index , level with GPT-6 Luna (max) and the top score from outside the US and China. It scores 50 on the Cyber Index and 82% on CyberGym-E2E-AA. Cost is $1.13 per task, over 4x that of similar-intelligence open models. Vals : It ranks #1 open-weight on HLAB and #9 among open models on the Vals Index . Heavy context use pushes its cost to $13.78 per test . Clinical triage : One evaluator reports a tie for #1 on 669 clinical decisions with zero severe misses. Caveats and disagreement : Refusal effect : Cline attributes the cyber lead largely to fewer refusals , saying Opus 5.5 and Astra had about 40% of tasks blocked by their own safety filters. Index gap : Critics note it trails GLM-5.3 and even GLM-5.3-Flash on AA’s index . Open-weight claim : Hugging Face’s CEO points out it isn’t open-weight until the weights ship . Configuration : Mistral warns that many reported failures come from not setting reasoning_effort="high" . Distillation hypothesis : Yuchen Jin speculates, as an unconfirmed opinion, that the Western–Chinese open-model gap reflects Chinese labs’ ability to distill Anthropic and OpenAI models . Open-Weight and API Model Releases: Embeddings, Image, Decision Models EmbeddingGemma 2 : Google’s first natively multimodal open embedding model covers text, code, image, video and audio in one space. It is built on Gemma 4 and released under Apache 2.0 ( DeepMind ). Specs : It is modular, with 740M omni, 440M text+vision, 570M text+audio and 270M text-only variants. It has Matryoshka dimensions from 768 down to 128, 8,192 context and a reported +14% on MTEB Code ( Phil Schmid ). Footprint : It uses roughly 191–567MB of active RAM and handles up to 5.5 minutes of audio or 58 video frames per pass ( Google ). Ecosystem : Day-0 support covers llama.cpp , vLLM , Ollama and Unsloth . It also runs in the browser on WebGPU at ~20–70ms per query . Nano Banana 2.1 : Google’s updated image model is rolling out across the Gemini app, AI Studio, Search and Ads ( Google ). Pricing : $0.034 per image, versus $0.134 for the previous Pro model, which Google says it outperforms ( Schmid ). Arena results : It ranks #4 in Multi-Image Edit, #5 in Text-to-Image and #6 in Image Edit, gaining +80 points over Nano Banana 2 in Text-to-Image ( Arena ). Decision models become a product category : OpenAI Decisions API : The public beta runs on GPT-6 Luna and returns predicates, choices or scores. OpenAI says it is up to 10x faster than the Responses API ( OpenAI Devs ). Pricing starts at $0.10/M input with no output charges . Perplexity : pplx-decider-v1.1-27b is open weights, costs $0.02/M input and tops the new HF Decision Index v0.3. Independent check on Jev : Vals found Jev matched GPT-6 Astra’s 97.5% on claim verification at about 1/500th the cost. Jev also ranked last on LegalBench . Skeptic view : Theo argues model-routing use cases are “absolutely useless” for choosing intelligence levels. Other open releases : Ling 3.1 Flash : The model has 560B total and 25B active parameters and scores 41 on AA’s index , up from 20. It costs $0.30/$0.90 per million tokens, and weights are coming. Reflection Beam : A Zhihu analysis of Beam describes a 501B/23B MoE with 23.8T pretraining tokens. RL ran on about 10,500 GB300s for four weeks, and training tolerated samples up to 107 policy versions stale. Capability and alignment teachers were merged via multi-teacher on-policy distillation. Kandinsky 6.0 : The video model ships under an MIT license with synchronized audio and day-0 vLLM-Omni support . Search eval : OpenAI’s built-in web search scores 74 on the AA Search Index , 5th among providers, at about $0.05 per task. It is weakest on BrowseComp, where it ranks 13th of 26. Safety, Control and Eval Integrity Control-intervention awareness :

FDE 判断

这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。