Headlines
01
Qwen ships Image 2.1
Alibaba's Qwen team releases version 2.1 of its open image-generation model family — the day's most-discussed AI launch at 672 points.
02
Google open-sources AX, an agentic orchestrator
AX is Google's open orchestrator for building and running agent systems — the day's biggest agentic infrastructure release.
03
Samsung to more than double HBM4 output next year
Samsung plans to more than double production of HBM4 and HBM4E DRAM next year, per supplier sources — the memory class feeding AI accelerators.
04
ZuckOff tells you when Meta's camera glasses are in the room
A free app that detects Meta smart glasses in the room — counter-surveillance tooling for wearable cameras, and the hottest story on the front page.
05
Where the Snowden archive went
An account of what became of the Snowden document archive — where it sits, and who can still read it.
AI / LLM / Agents
01
Kev: a tiny Jev-like family of decision models on Qwen3.5
Jared Palmer's Kev builds a small family of decision models on top of Qwen3.5 — models that decide rather than generate.
02
Mini-AGI: continual learning on 8GB of VRAM
A dynamic continual-learning model trainable on 8GB of VRAM — weight updates that keep going on consumer hardware.
03
The case that MCP was always a bad idea
An essay arguing the Model Context Protocol was a mistake from the start — the debate fills a 161-comment thread.
04
Lossless-memory: personal AI memory without summarisation
Hours old on the front page: a personal memory layer for AI that stores context verbatim instead of summarising it — no lossy compression of your history.
05
Heretic strips restrictions from language models
A project that removes restrictions from language models — automated removal of refusal behaviour from open weights.
Papers — ArXiv CS.AI
01
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
RL for coding agents is bottlenecked by task supply — existing pipelines extract environments from development artefacts such as issues and commits, capping volume and diversity. CodeMidas takes source code alone as input: agents explore implemented functionality to write behavioural specifications, build tests grounded in execution of the original code, then filter candidates via repeated solution rollouts. The output is 5,545 training tasks from more than 3,100 codebases, with agentic compute spent at every stage of environment construction. [agentic coding, reinforcement learning, synthetic data]
02
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
Full benchmark reruns on every revision of a production agent are too expensive, yet most evaluation research assumes one-off academic runs. Across 574 historical runs of an analytics agent serving tens of thousands of monthly users, multidimensional 2PL adaptive testing gave the best fidelity — 200 questions, 38.5% of a full run, held MAE to 1.03 pp. The authors deployed the simpler difficulty-stratified fixed subsets instead, and show those transfer without recalibration to five other agent families. [evaluation, adaptive testing, production agents]
03
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
Sparse prefill schemes that rank blocks by centroid similarity miss the case where one highly relevant token hides inside a mediocre block — the authors name this failure mode mean dilution. RBS-Attention pairs the centroid branch with a rescue branch driven by maximum key-block radius, thresholded independently so execution still runs as regular block-sparse FlashAttention. Training-free, it delivers a 20.65× speedup on standalone prefill attention with H100 GPUs. [sparse attention, long context, inference efficiency]
04
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Long-horizon agentic work such as professional graphic design admits no programmatic oracle, so improvement normally stops when the weights freeze. Designer-RSI moves the learning outside the model — a frozen frontier model drives more than 230 design tools while an external bank of natural-language procedures widens with new subtasks and deepens as procedures are revised against their own recorded successes and failures. A matched replay gate admits only changes that repair failures without regressing past successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories grow the bank from 76 skills, with no weight updates and no human labels. [procedural memory, agentic design, continual learning]
05
GUARD: Natural Forgetting in Large Reasoning Models
Unlearning is harder in reasoning models — a protected fact can surface in chain-of-thought before the final answer, and suppression objectives never specify what the post-forgetting trace should contain, inviting hallucinated substitutes or repetitive output. GUARD converts model-generated unsafe disclosures into safe-exit trajectories — coherent non-disclosing reasoning followed by a stable refusal — and distils them into a frozen model via guidance tokens. The position: forgetting needs a shaped successor behaviour, not a hole. [unlearning, reasoning models, safety]
06
HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
Homomorphic encryption lets a server evaluate an LLM over a client's encrypted prompt — which also blinds it to jailbreaks, so a malicious client's attack is invisible by construction. HE-Guardrail evaluates guardrail mechanisms directly over encrypted inputs, restoring adversarial screening without surrendering client confidentiality. Privacy-preserving inference has mostly been demonstrated on benign clients — this addresses the hostile one. [homomorphic encryption, jailbreak defence, privacy]
Hacker News
01
Grim Fandango's 1996 puzzle design document
The original 1996 puzzle design document for Grim Fandango, scanned to PDF — adventure-game design from the LucasArts era.
02
What CRTs actually did to pixel art
Why pixel art was drawn for CRTs — scanlines and phosphor glow were part of the medium, and flat panels took them away.
03
Amiga Unix, again
A revival of Amiga Unix — Commodore's short-lived System V Release 4 port for the Amiga.
04
The senior engineer death spiral
Sunil Pai on the self-reinforcing failure mode that catches senior engineers.
05
Sherline Tools is shutting down US production
Sherline, the long-running American maker of desktop lathes and milling machines, is going out of business.
Why It Matters
▸
The agent stack is commoditising from the top down. Google open-sources AX and CodeMidas mints 5,545 verifiable RL environments from raw source with no human task authorship, which makes the 161-comment case against MCP read like a dispute over a layer the field is routing around. When orchestration is free and environments are generable, the scarce goods move to verifiers and recurring evaluation — the production benchmarking study is the honest companion piece, and its authors' own choice is the lesson: adaptive testing reached 1.03 pp MAE at 38.5% of full-run cost, and they still shipped the simple fixed subset because operations beat optimality. Qwen Image 2.1 led the day's discussion at 672 points — the durable releases were plumbing.
▸
Adaptation pressure is being answered two ways — buy memory or change its shape. Samsung will more than double HBM4 output next year, the capital answer, while RBS-Attention's 20.65× prefill speedup and Designer-RSI's growing skill bank are the software answer, extracting more from hardware already bought. The frozen-weights pattern runs down the whole stack — Kev's tiny decision models that decide rather than generate, Lossless-memory keeping history verbatim instead of summarising it, Mini-AGI squeezing continual updates into 8GB of VRAM. The counter-image is Sherline ending US production — memory silicon compounds, machine tools don't.
▸
The common premise of this week's safety artefacts is that upstream control has already failed. Heretic strips refusals from open weights automatically, so alignment can no longer be assumed from a checkpoint. GUARD handles the reasoning-model variant — a protected fact resurfaces in chain-of-thought before the refusal lands, so it distils a shaped successor trajectory rather than a hole — while HE-Guardrail covers the opposite blindness, a server running encrypted inference that cannot read the prompt it is guarding. ZuckOff is the same instinct in hardware, detecting the camera you cannot inspect, and the Snowden archive piece is its decades-old precedent — control over information is a property of who can still read it.