AI and LLM Weekly — 10 October 2026
OpenAI releases a large batch of AI-produced mathematics, while Anthropic ships a far cheaper Haiku 5.5 and Mistral previews a trillion-parameter open-weight model.
OpenAI publishes a large repository of mathematical results from an unreleased model
OpenAI · 6 Oct 2026 · Research
OpenAI put a collection of new mathematical results, generated by an internal frontier model it has not named or released, into a public GitHub repository. Many of the proofs come with Lean formalizations that a computer can check, and OpenAI says the average result took roughly three hours of ChatGPT Pro-equivalent compute. Not every proof is formalized yet and none have been through journal review. The release follows the company’s setup of a mathematician advisory group, whose guidance it says shaped how the results were shared.
Claude Haiku 5.5 arrives at about a tenth of its predecessor’s input price
Anthropic · 7 Oct 2026 · Model release
Anthropic’s new small model is priced at $0.10 per million input tokens and $0.50 for output on prompts up to 100K tokens, and it is the first Haiku with an adjustable effort setting. Reported scores include 72.4% on OSWorld 2.1 and 39.2% on Terminal-Bench 4.0, well ahead of Haiku 4.5 though behind Sonnet 5.5. Anthropic also halved Sonnet 5.5 cache-read pricing. It matters for high-volume work such as classification and subagent roles.
Mistral previews Large 4, a roughly trillion-parameter open-weight mixture-of-experts model
Mistral AI · 6 Oct 2026 · Model release
Mistral Large 4 combines instruction-following and reasoning in one multimodal mixture-of-experts model with about 1 trillion total and 52 billion active parameters. A preview API is live at $1.36 per million input tokens and $4.18 for output, with weights promised by the end of October. Mistral’s own coding and agentic benchmarks put it near the top frontier models, though these are self-reported. The open release is the part to watch.
Anthropic launches a long-term Cyber Mission with programs for infrastructure and open source
Anthropic · 8 Oct 2026 · Security
The initiative starts with a Critical Infrastructure Defense Program, which gives operators of power, water and transport systems frontier models, on-site engineers and threat research, with firms such as CrowdStrike and Palo Alto Networks as founding partners. It also introduces a free, opt-in OSS Scanner that audits enrolled open-source projects and reports a proof of concept and suggested fix. Anthropic expects a true-positive rate above 90% but notes the reports are not human-reviewed.
Reflection AI unveils Beam, a 501-billion-parameter open-weight coding model
Reflection AI · 5 Oct 2026 · Model release
Beam is a sparse mixture-of-experts model with 23 billion active parameters, built with a 1M-token context in mind and trained on 23.8 trillion tokens. Reflection reports 80.9 on SWE-bench Verified, though it trails several larger rivals on harder coding tests. Access is limited to a waitlist for now, with Apache 2.0 weights and a technical report promised later in October.
ChatGPT’s GPT-6 answers can now come as interactive interfaces
OpenAI · 7 Oct 2026 · Product
A new Intelligent UI feature lets GPT-6 reply with buttons, forms, charts and small working tools such as calculators instead of plain text, choosing the format itself. Paid plans get it first, free and Go users a day later, with Sol and Luna as the underlying models. OpenAI also says web-search answers begin 44% sooner than with GPT-5.6 Instant.
Anthropic rewrites its Usage Policy, with new rules on surveillance and autonomous hardware
Anthropic · 8 Oct 2026 · Policy
The updated policy takes effect on 12 November 2026. It bars using Claude to decide who police should investigate or charge, requires a qualified operator to be able to stop equipment that Claude controls, and adds a narrow ban on sustained, purposeless cruelty toward the model. The blanket ban on personalized campaign targeting is dropped, though deceptive targeting stays prohibited.