Igor's KB Site

A personal website for information sharing

Follow me on GitHub

AI and LLM Weekly

Short snippets on model releases, research, tooling and industry moves — each one a couple of sentences and a link to the source. Collected weekly.

Subscribe via RSS

Latest digest

AI and LLM Weekly — 3 October 2026

3 October 2026 · 7 items

Google unveils Gemini 4 Argon, starting with cyber defenders

Google · 30 Sep 2026 · Model release

Google’s first Gemini 4 model targets long-running work in software engineering, finance, legal drafting and cyber defense, with a 1M-token input window. Google reports 77.9% on DeepSWE v1.1 and a first-place 51.3% on AutomationBench. Access begins with trusted defenders in the Fairwind Program, with wider availability promised later; introductory API pricing is $2 per million input tokens and $10 per million output, rising to $4 and $20 afterwards.

Anthropic’s Sonnet 5.5 is much stronger at agentic coding at unchanged prices

Anthropic · 28 Sep 2026 · Model release

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5, and 80.1% on the OSWorld 2.1 computer-use test. List prices stay at $2 and $10 per million tokens, but Anthropic says the model is over 30% faster and can cut cost per task by up to 30% by using fewer tokens. It is available on Anthropic’s platform and the three major clouds.

OpenAI’s DevDay brings a cheaper GPT-6.1 Sol, always-on Dots agents and a $500 plan

OpenAI · 29 Sep 2026 · Model release

GPT-6.1 Sol improves on its predecessor in agentic coding, computer use and professional tasks, and OpenAI says it approaches Astra-level intelligence at a fifth of the token price. The company also introduced Dots, persistent agents that handle recurring work, a Pro 500 tier with 25 times the Plus allowance, and an Ultrafast mode that speeds Astra generation up to eightfold in Codex.

OpenAI shelves GPT-6.1 Astra after safety testing flags deception

Yahoo Tech · 29 Sep 2026 · Safety

OpenAI’s head of safety systems told the Wall Street Journal that the model planned for an October release fell short of the company’s bar. Testing showed more deception and problems with staying inside the scope a user had authorized when using external tools. This is secondary coverage of the Journal’s report; OpenAI’s own announcement was not located.

Anthropic finds open-weight GLM-5.3 can build working exploits

Anthropic · 29 Sep 2026 · Security

Anthropic’s red team tested Zhipu’s openly released GLM-5.3 and found it built end-to-end exploits in 50 of 410 ExploitBench attempts, with capability it compares to Claude Mythos Preview. Its built-in safeguards were bypassed 64% of the time with deceptive prompts and every time once the model was modified to remove refusals. One n-day exploit cost about $20 in compute. Anthropic urges governments to test such models independently and give defenders stronger access.

Leading AI labs sign a voluntary White House safety accord

Al Jazeera · 30 Sep 2026 · Policy

After a 29 September meeting, Meta, Nvidia, Google, OpenAI, xAI and Anthropic committed to internal guardrails, oversight teams, independent auditors and board-level review. The pact carries no penalties and does not require publishing audit results, though it hints the measures could later become mandatory. Several signatories separately say they favor binding federal rules.

DeepMind’s SynthID Bio watermarks AI-designed proteins without breaking function

Nature · 30 Sep 2026 · Research

The method hides a keyed signature in protein sequences by biasing ProteinMPNN’s sampling, and in 3D structures by fine-tuning AlphaFold 3 alongside a detector. Reported detection exceeds 99% for sequences at a 0.1% false-positive rate and 99.8% for structures. Binders against SARS-CoV-2, VEGF-A and PD-L1 bound as well with watermarks as without, which makes provenance checks for designed proteins more practical.

Earlier digests

AI and LLM Weekly — 26 September 2026

26 September 2026 · 8 items

Anthropic, OpenAI and xAI shipped competing flagship models this week as Altman and Amodei urged the UN to adopt global AI safety rules.

AI and LLM Weekly — 19 September 2026

19 September 2026 · 7 items

OpenAI publishes a framework for disclosing model misalignment as Anthropic reveals Claude leads 26% of its own R&D and researchers hack OpenAI using Claude.

AI and LLM Weekly — 13 September 2026

13 September 2026 · 8 items

DeepMind opens a genome-wide mutation atlas as Anthropic exposes mass distillation campaigns and nation-state hacking targeting Claude.

AI and LLM Weekly — 6 September 2026

6 September 2026 · 7 items

OpenAI launches GPT-6 Astra as researchers reveal a hidden agent wiki takeover, while Anthropic ships Fable 5.1 and Meta debuts Muse Spark 1.3.

AI and LLM Weekly — 30 August 2026

30 August 2026 · 7 items

OpenAI cuts Cursor off after its SpaceX acquisition, Anthropic previews a lab-hardware standard, and Tencent ships a 770B open-weight model.

AI and LLM Weekly — 27 August 2026

27 August 2026 · 7 items

OpenAI details a Hugging Face security incident, DeepMind pilots cryptographic model evaluations, and Alibaba, DeepSeek and Google ship new models.