← Index
AI NEWS DAILY

The AI Briefing

Non-autoregressive RL models resurface · Microsoft director calls AI scraping the largest theft of labour · StepFun previews Step 5 · Spain blocks Archive.today
AI NEWS DAILY · 20 SEP 2026 · EN 0:00 / 5:00
01
Non-autoregressive decision models with RL top HN
A builder's retrospective on non-autoregressive decision models trained with RL, posted a year after the work. 1,238 points, 295 comments.
02
Microsoft director calls AI scraping the largest theft of labour
Legal briefs in the NYT lawsuit quote a Microsoft director calling AI scraping the largest theft of labour in human history, while OpenAI's head calls ChatGPT an existential threat to publishers. 171 points, 47 comments.
03
Exfiltrate Your Weights
A site on exfiltrating model weights, near the top of HN with 470 points and 186 comments.
04
RSA-896
A technical post on RSA-896, the 896-bit RSA modulus. 163 points, 60 comments.
05
Spain orders blocks on Archive.today and mirrors
Spain has ordered blocks on Archive.today and its mirrors, drawing 173 points and 164 comments.
01
StepFun previews Step 5
StepFun's Step 5 preview claims advances on the Pareto frontier. 85 points, 23 comments.
02
Cache-to-Cache: direct semantic communication between LLMs
A 2025 paper on direct semantic communication between LLMs through cache-to-cache transfer. 105 points, 19 comments.
03
Stagehand makes Playwright 2x faster and 80% more token-efficient
Browserbase's Stagehand reports a 2x Playwright speed-up and 80% token reduction for its browser automation stack. 140 points, 36 comments.
04
Terry Tao asks whether we need human mathematicians anymore
Terry Tao's essay on what human mathematicians are for as AI systems improve. Posted near the top of HN.
05
TMLR asks authors about their own papers
A TMLR post on asking authors about their own papers, drawing 179 points and 96 comments.
06
AI writing is almost never worth it
An argument that AI should almost never be used to write, drawing 308 points and 150 comments.
01
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with RL see one scalar reward per trajectory — sparse signal for long-horizon credit assignment. On-policy distillation from a privileged self-teacher promises dense token-level supervision, but privileged information does not guarantee a reliable teacher, and its value is stage-dependent. RetireOPD trains a skill-conditioned teacher on environment rewards first, then trains a skill-free student on RL plus distillation, with Adaptive Retirement: the student drops the teacher once their discrepancy stops shrinking. reinforcement learning, distillation, agents
02
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Support-agent RAG treats closed cases as static documents, which fails when the useful signal is a case's intermediate states rather than its conclusion. RAFT abstracts each historical case into a directed chain of timeline entries, retrieves at entry level, and returns the parent-case trajectory anchored at the matched state — with an optional case-level similarity graph on top. Because the retrieval layer is evaluated directly, results need no production deployment. retrieval, agents, enterprise support
03
MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation
LLM causal discovery rarely treats premature omission of potentially relevant relations as an explicit design objective, so candidate graphs get pruned before direct-edge judgement ever runs. MaSCoD organises candidate third variables and local structural patterns first, then judges direct edges. On Auto-MPG, DWD and Sachs — GPT-5.4 primary, GPT-4o for replication — supplying structural hypotheses before judgement beat the no-phase-one ablation on mean recall and F1 across all six dataset-backbone settings, while raising false-positive rates. causal discovery, multi-agent, evaluation
01
Btrfs, ZFS and bcachefs under workloads classic benchmarks skip
A filesystem benchmark comparing Btrfs, ZFS and bcachefs on workloads that standard benchmarks miss. 144 points, 133 comments.
02
Tin: full-text search for Postgres
PlanetScale introduces Tin, full-text search for Postgres. 218 points, 86 comments.
03
OONI: run internet-censorship measurements
OONI's install page for measuring internet censorship from your own network. 169 points, 108 comments.
01
E-ink frame hears birds and draws them as 1800s illustrations
A Show HN e-ink frame that listens for birds and renders them as 19th-century illustrations. 2,366 points, 259 comments.
02
English: a vs an
A deep dive into English indefinite article selection. 270 points, 345 comments.
03
What Zig felt like coming from Rust
A Rust developer's experience moving to Zig. 235 points, 286 comments.
04
Defeating Chrono Trigger's Dream Devourer with an int overflow
A write-up on using an integer overflow to beat the Dream Devourer in Chrono Trigger. 136 points, 77 comments.
05
How SpaceX streamlined the Raptor engine
Construction Physics on how SpaceX streamlined the Raptor engine. 238 points, 107 comments.