Headlines
01
Claude Fable 5.1 and Claude Mythos 5.1
Anthropic announces Claude Fable 5.1 and Claude Mythos 5.1, the day's most-voted story by a wide margin.
02
Hang on to Your Firefox
Newsonaut makes the case for sticking with Firefox as browser competition and default search deals come under pressure.
03
AnkiDroid: Google Play no longer allowing Open Collective donation link
Google Play is rejecting AnkiDroid's Open Collective donation link, threatening one of the largest open-source Android projects' funding route.
AI / LLM / Agents
01
How accurate have Ed Zitron's AI skeptic predictions been?
Dan Luu checks Ed Zitron's past AI-skeptic predictions against what actually happened since he made them.
02
I trained a small transformer in 1.5hrs and it beats many LLMs
The author trains a small transformer from scratch in about 90 minutes for the ARC benchmark and reports it beating many general-purpose LLMs.
03
The ChatGPT/Codex app bundles a full copy of LibreOffice
Simon Willison digs into the Codex desktop app and finds it ships an entire LibreOffice install, raising questions about desktop-app weight.
Papers — ArXiv CS.AI
01
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
First pragmatic benchmark for Hindi, Punjabi, Tamil and Malayalam, covering deixis, speech acts, implicature, social pragmatics and coherence. Multilingual LLMs fail consistently on meanings rooted in Indic conventions, and translation quality does not reliably track pragmatic understanding. [eval, multilingual, pragmatics]
02
Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
Tool Primitives swap rigid API schemas for natural-language tool interfaces, and ToolFace retrieves from a repository of 25,519 functions at inference time instead of enumerating schemas in context. The HEART planner/router/verifier harness beats SFT baselines by 10% and frontier commercial models by 6% on average while cutting API cost up to 85%. [agents, tool-use, inference]
03
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
A two-stage recipe: a 4B teacher reranker is strengthened with off-policy GRPO on LLM-judge feedback, then a 1B student is distilled on-policy over rankings it samples itself. The student reaches 0.7670 nDCG@6 on MAIR-11, beating offline listwise KD by 4.6 points, with gains concentrating under distribution shift. [distillation, reinforcement-learning, reranking]
04
Thinking effort aligns between humans and reasoning models in abductive reasoning
Uses abductive reasoning, whose difficulty cannot be read off formal structure, to test whether reasoning models' thinking effort tracks human reaction times without shortcuts. Finds further alignment in reasoning cost and error patterns, and that decoding methods exploring multiple paths increase the correspondence across the three models tested. [reasoning, eval, cognitive]
05
Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks
Treats inter-agent messages as traces of a sender's reasoning and models 6G RAN agents as cognitive channels on a cellular sheaf, so protocol-valid reports can still carry a hallucination. Derives five design principles for resilient multi-agent systems and validates them on locally deployed 1B-parameter telecom models, where a cognitive SNR isolates a hallucinating peer that three of four neighbors agree with. [agents, multi-agent, reliability]
Why It Matters
▸
Model releases and model criticism are pulling in opposite directions on the same front page: the day's top post is a 5.1 bump from Anthropic, while the fourth is an audit of an AI skeptic's track record. Both are really asking how much of the promised capability has shipped.
▸
Today's ArXiv batch is mostly about measurement — a pragmatic benchmark for Indic languages, a 1B reranker distilled to hard nDCG numbers, and a study of thinking effort in abductive reasoning. Even the 6G paper is about detecting the peer a model cannot trust, not about raw capability.