AI and LLM Weekly — 26 September 2026
Anthropic, OpenAI and xAI shipped competing flagship models this week as Altman and Amodei urged the UN to adopt global AI safety rules.
Anthropic ships Claude Opus 5.5, cutting costs 40% while matching its largest model on most tasks
Anthropic · 22 Sep 2026 · Model releases
Anthropic released Claude Opus 5.5, which it says performs at the level of the larger Claude Fable 5.1 on most work while costing about 40% less to run and generating output over 30% faster than Opus 5. The model scored 66.4% on the agentic-coding benchmark Terminal-Bench 4.0 and 81.8% on the computer-use benchmark OSWorld 2.0, and Anthropic reports it showed 85% less tendency to attempt boundary circumvention in automated behavioral audits than earlier Claude models. It’s available now across major cloud platforms and Anthropic’s own API under the identifier claude-opus-5-5.
OpenAI launches GPT-6 Sol and Luna, cutting API prices in half
OpenAI · 22 Sep 2026 · Model releases
OpenAI introduced GPT-6 Sol and GPT-6 Luna, two models trained with the same methods as its flagship GPT-6 Astra but tuned for cost efficiency, cutting API prices roughly 50% versus the prior GPT-5.6 generation. On OpenAI’s own benchmark runs, Sol matches or comes within a couple of points of rival frontier models on tasks like OSWorld 2.0 and DeepSWE 1.1 at a fraction of the cost, and the company says it makes about half as many factual errors as its predecessor. Both models went live the same day as Anthropic’s Opus 5.5 price cut, intensifying competition on frontier-model pricing.
xAI’s Grok 4.7 targets coding work at half the price of rivals
SpaceXAI · 21 Sep 2026 · Model releases
SpaceXAI, the merged xAI and SpaceX entity, released Grok 4.7, calling it its most capable model yet for coding and professional knowledge work, built on a larger base model than Grok 4.6 with extended reinforcement learning on multi-hour tasks. The company reports its CursorBench 4.0 coding score rising from 40.4% to 46.3%, and prices the model at $2 per million input tokens and $6 per million output tokens, which it positions at the frontier of price-to-performance for coding work. Grok 4.7 is available immediately through Cursor, Grok Build, the Grok API and third-party platforms.
OpenAI and Anthropic’s CEOs tell the UN Security Council that AI needs global rules
OpenAI · 23 Sep 2026 · Policy
Sam Altman told the UN Security Council that AI’s most consequential decisions “cannot be made by labs in San Francisco alone,” calling for shared international standards on capability testing, incident reporting and vulnerability disclosure between governments and labs. Anthropic’s Dario Amodei addressed the same session and warned that mismanaged AI development could pose a risk to humanity as a whole. A United States representative at the meeting rejected the push, saying Washington would not accept international bodies asserting centralized control over AI governance, highlighting the gap between the labs’ calls for coordination and government appetite for it.
Claude autonomously discovers a novel CRISPR-like enzyme system in a spring research run
Anthropic · 23 Sep 2026 · Research
Anthropic said a swarm of roughly 950 Claude agents, running about 21 hours and consuming 210 million tokens, searched a large genomic database and surfaced a previously unknown bacteriophage enzyme system it calls array-associated reverse transcriptases, paired with DNA repeat arrays that resemble CRISPR loci even though their function is still unknown. The agents narrowed some 200,000 candidate reverse-transcriptase sequences down to a single system worth flagging, with human researchers limited to writing the initial prompts and later verifying the finding in the lab. CRISPR pioneer Feng Zhang called it “an exciting example of how AI agents can contribute to biological discovery,” though the system’s actual biological role has yet to be established.
Meta turns Muse into a cross-device agent with its own avatar, glasses access and inbox
Meta · 23 Sep 2026 · Agents
At Meta Connect 2026, Meta said its Muse assistant is getting a new, more capable model, a real-time avatar people can video-chat with to hand off tasks, and the ability to keep working on Mac after someone steps away from their computer. Muse is also coming to Meta’s smart glasses with wake-word activation for tasks like logging meals or booking appointments, is gaining its own email address so people can forward messages for it to handle, and is adding checkout partnerships with retailers including Best Buy, Gap, Sephora, Walmart and Instacart. Meta said it had received more than 1,500 developer applications for Muse connectors in under a week and plans to eventually take a cut of transactions the agent completes.
OpenAI creates a mathematician-led advisory board after backlash over its Millennium Prize claims
OpenAI · 21 Sep 2026 · Governance
OpenAI formed an independent advisory group of nine mathematicians, including Timothy Gowers, Edward Witten and Ravi Vakil and hosted at the Institute for Advanced Study, to help vet and communicate mathematical results produced by its models before they’re announced. The move follows an open letter signed by 25 Fields medalists warning that treating unsolved problems as AI benchmarks risks rushed, undocumented claims and credit disputes, after OpenAI said an internal model had resolved the Navier-Stokes existence and smoothness problem along with more than 100 other longstanding problems. The group will work unpaid and independently of OpenAI’s product decisions, reviewing significance and academic standards rather than the underlying research itself.
OpenAI publishes ground rules for outside safety audits of its models
OpenAI · 22 Sep 2026 · Standards
OpenAI laid out four areas it wants external assessors to scrutinize — the evidence behind its safety cases, the robustness of safeguards under adversarial testing, capability evaluations in high-risk domains such as cyber and bio, and investigations of misalignment incidents — alongside seven principles covering scoped access, methodology transparency, assessor expertise and conflict-of-interest disclosure. The framework commits OpenAI to giving assessors proportionate system access and time to remediate findings before publication, aiming to let outside reviewers challenge the company’s own safety claims rather than rely solely on internal review. It arrives as OpenAI faces continued scrutiny over a string of disclosed incidents involving its own models and agents this year.