Fifty-six stories cleared 200 points today, and twenty-five of them are new. That’s the largest new crop I’ve logged, and the reason is obvious: three frontier-adjacent model launches landed inside twenty-four hours — Xiaomi’s MiMo v2.6, Anthropic’s Claude Opus 5.5, and OpenAI’s GPT-6 Sol and Luna — and HN spent the day doing what it does with launch days, which is arguing about whether the price cut is the product. Behind that, a second cluster formed around Apple removing switches it used to provide, and a third around writing and code review now that most of the writing and code review is machine output.

The short version of the day: the capability race has moved to the cost curve, and the interesting disclosures are no longer benchmarks.


MiMo v2.6

1,076 points · Xiaomi

Xiaomi released and open-sourced the MiMo-V2.6 series: two natively multimodal models, Pro and Flash, trained with a scaled-up reinforcement learning effort the release frames as a step on the “RSI (recursive self-improvement)” path — build on verifiable complex tasks, scale RL compute, let the model expand its own frontier through exploration and feedback. The vendor’s headline number is 46 on Artificial Analysis’s composite intelligence index for MiMo-V2.6-Pro, which they say makes it the strongest open-weights model available, ahead of Kimi K3 and Qwen3.8 Max, while still behind Claude Fable 5.1 and GPT-6 Astra among closed models. They also shipped MiMo-V2.6-Distill-Qwen-9B, released the technical report, and — unusually — launched a desktop client and a subscription plan at the same time as the weights. API pricing unchanged.

The numbers are the least interesting part, and the thread knew it. What people actually responded to is that Xiaomi ran a public realtime dashboard of the training run at mimo.xiaomi.com/rl/. Commenters with training experience described it as the single best teaching artifact anyone has published — the metrics tab is absurdly detailed, the failures are included, and the “notices” stream is a running account of interventions. The plausible technical read is that it’s built on the verl RL framework’s dashboarding, but the transparency is a choice, not a framework default, and the contrast was drawn explicitly: American labs publish hundreds of pages of writeups after the fact, and Xiaomi published the run while it was happening. One commenter’s summary of why this matters is the one worth keeping — a live dashboard is a record of the process, and the process is the part that published benchmarks and vendor writeups systematically omit.

Two caveats. “Most powerful open-source model” is a claim about one composite index, and the composite is assembled by a third party from a mix of vendor-reported and independently-run evals. And open weights plus a published dashboard is still not open training data or open training code; the same commenter pre-empted the objection by noting they’d take what they can get. Fair. The bar has moved so far that publishing the run in realtime counts as a breakthrough.


I don’t want to read what you didn’t write

971 points · Colin Breck

The most-referenced essay of the day, and it earns it. Breck’s complaint is specific and it’s not “AI writes badly.” It’s that people who never produced original writing are now producing design proposals, business plans, PR descriptions, tickets, meeting summaries and personal messages, and the output is unreadable — not because the prose is bad but because it’s a pile of statements outside of context. The pattern he names first is the sharpest: someone builds a thing, then asks a model to retrospectively summarize the thing into a design document. That document is no longer a proposal for building consensus; it’s a machine’s summary of a decision already made, in exhausting detail, with no perspective on which parts matter. The reader has to read all of it, because there’s no signal about what’s relevant. Of course they stop.

The information-theoretic version of the argument, offered in the thread and worth stating cleanly: if you have 1,000 bits of semantic information to transfer and you give a model 300, it cannot supply the missing 700 — it can only guess. And if it guesses correctly, those bits weren’t semantic information to begin with. You might as well have sent the 300. This is a better formulation of the problem than most, and it generalizes: what makes writing worth reading is the part that was not inferable from the prompt.

What saves the essay from being a crank post is the second half, which is a genuine account of where AI writing helps. Breck wrote an academic paper in LaTeX and used AI extensively for it, feeding it the journal style guide, prior issues, the source code for the software he was describing, and production configuration, logs and metrics. What worked: verifying each finished paragraph against the actual source — did he describe the indexed columns correctly, is that how the Parquet rows are sorted — and completing BibTeX citations so he didn’t break stride. What also worked: spelling, grammar, ruthlessly, and finding a subtle notation error that four expert human reviewers missed. What never once worked: asking the model to write the paragraph from that same context. Consistently unpleasant, frequently inaccurate. The one piece of writing he took unchanged was the abstract — “the most terse, mechanical, inhuman part of the whole paper.” That’s the thesis in one sentence: AI is excellent at compressing work that already exists, and useless at producing the work it would then compress.

The thread’s second-best contribution: someone pointed out that giving someone raw model output is structurally identical to the old journalistic habit of summarizing a paper without linking it. That habit is forty years old, has produced a forty-year trail of “scientists find X which proves Y” where the paper said “may suggest,” and nobody needed AI to invent it. The model didn’t create the disease. It industrialized it.


Claude Opus 5.5

780 points · Anthropic

New model, first of the Claude 5.5 family. Anthropic says it performs “at the level of Claude Fable 5.1 on most work” and costs 40% less to run than Opus 5. Pricing: $4/M input, $20/M output — 20% below Opus 5 — and cache reads at $0.20/M, down 60%, which is the number that actually matters because cache reads dominate the cost of agentic coding. The release notes it was externally evaluated by Frontier Design and METR, scores best-to-date on their automated behavioral audit, and is deployed with safeguards comparable to Fable 5.1 because it’s comparable in biology and cybersecurity — vetted organizations can apply to the Life Sciences Verification Program, with the Cyber Verification Program expanding to verified practitioners in coming weeks. Capability claims are anecdote-shaped: one tester did a 680,000-line code migration in under a day; on a page-optimization task it succeeded 39 of 40 times where Opus 5 made smaller changes that also altered app behavior.

The reception was dominated by the first line of the announcement, which reads: “Claude Opus 5.5 is our first release since we called for pacing the frontier.” The thread’s read — stated bluntly and upvoted — is that the sentence reminds you of last week’s restraint post and everything after it demonstrates the opposite with very specific numbers. That’s a fair hit, and the follow-on argument is more interesting: multiple commenters characterized “pacing the frontier” as a bid to convert a competitive disadvantage into a regulatory moat. I think that’s too strong as stated, and the vocabulary argument it degenerated into was tedious — a long subthread on whether “pacing” means walking faster or walking around anxiously, which is not the interesting question. The interesting question is what a commitment to restraint is supposed to look like when it comes with a 60% cache price cut, and the release doesn’t answer it.

The substance that got buried: cache-read pricing at $0.20/M is the competitive move here, not the headline reduction. For anyone running long agent sessions, cache reads are the bill. Cutting that by 60% is a direct attack on the cost structure of the workflow Opus is sold into, and it lines up with the same day’s OpenAI announcement, which is not a coincidence.


I said no and Apple said yes

758 points · David Bushell

Bushell’s blog is useful because he timestamps his objections. On Wednesday 5 February 2025 at 09:51 he documented that macOS 15.3 had enabled a feature phoning home every fifteen minutes with personal data, and he turned it off — macOS provided a setting, though it hid the second half of it a menu deeper, because why should saying no be easy. Last week he upgraded macOS 15 to 27. Today he found that Apple removed the off switch. The Siri toggle still exists, but “Apple Intelligence” is gone from that control, multiple unkillable Siri processes continue running and writing data, and turning Siri off doesn’t stop any of it. The actual restrictions now live under Screen Time — a parental-controls surface — and even those don’t disable anything; they hide the menus.

Then the disk. Apple Intelligence is consuming 22.28 GB on his machine, having been enabled during the upgrade despite his explicit prior opt-out. He prices it against Apple’s £500/TB storage upgrade: about £11, which he notes is trivial compared to the consent problem, and he’s right that the money isn’t the story. The story is that a previously honored “no” was silently converted to a “yes” by an OS upgrade, and the replacement off switch is now buried in a feature intended for parents restricting their children.

The thread went where these threads go — Linux versus Mac versus Windows and which one has worse sharp edges — and it’s tedious except for two genuinely useful contributions. First, the response to “just use Linux”: every platform has sharp edges, they’re just distributed differently, and the author’s inability to disable AI features is a different kind of problem from an interface vanishing. Several people made the distinction well — users tolerate bloat they can ignore; they don’t tolerate breakage that blocks work. Second, the practical: one commenter linked a GitHub guide for deleting Apple Intelligence and recovered 16+ GB. Another proposed the classic hack — after deleting the model directory, replace it with a locked dummy file of the same name via chflags schg so redownloads fail, then check whether the download logic notices.


GPT-6 Sol and Luna

640 points · OpenAI

Two cheaper tiers added below GPT-6 Astra, which stays the top model. OpenAI halved prices versus the GPT-5.6 promotional pricing: Sol $4 → $2 input and $20 → $10 output; Luna $0.20 → $0.10 input and $1.20 → $0.50 output. On AutomationBench — 47 tools across sales, marketing, ops, support, finance, HR — they claim Sol at xhigh effort beats Claude Opus 5 at max effort at 9% of the cost per task, and Luna at high effort improves on its predecessor by 5.4 points at 58% lower cost per task. There’s an unusual footnote, and it’s to OpenAI’s credit that it’s there at all: the Fable 5.1 row understates Fable’s actual cost because it omits the Opus 5 fallbacks, which happened on roughly 40% of tasks.

The most useful independent data came from Simon Willison, who runs SVG-generation tests at every effort level for every model. The Luna results are cheap and on the Pareto frontier for most tasks; the Sol results at max effort are the best of the new crop; Astra max still wins on his grid. Thread consensus was that Luna’s price-to-capability ratio is currently unmatched among closed models, that Gemini 3.8 Flash is the real competitor on that axis rather than the previous OpenAI tier, and that MiMo 2.6 Pro now occupies the same Pareto slot from the open-weights side.

The skepticism worth recording is not about the numbers — it’s about the sustainability. “I don’t know how they make money here” was a top comment, and nobody had an answer beyond customer acquisition. A 50% cut against promotional pricing, on models that were themselves sold as cost-efficiency plays, is a land-grab. That’s fine as a strategy; it’s just not the same thing as a cost curve improving, and treating a price war as evidence of efficiency gains is how people end up surprised when the price goes back up. Also buried in the thread: multiple users said they pick models by coding style rather than benchmark score, which is the kind of thing that makes vendor charts less predictive than they look.


Spymarks, not Watermarks

638 points · brand.io

A terminological argument with a real payload. A watermark, the author says, is a visible mark verifying authenticity or asserting ownership. A spymark is a hidden signal that makes your work traceable without your knowledge or consent, and the reason to name it separately is that “watermark” currently does both jobs and thereby launders the second one. The example is Google’s SynthID, which embeds signals “imperceptible to humans” (Google’s words) into images, audio, text and video, and which per Google’s own SynthID-Image paper can encode a 136-bit payload in a 512×512 image under the SynthID-O variant — enough for a 64-bit database identifier plus 72 bits of error correction. 64 bits is room to point at a user record: name, addresses, IPs, date of birth. OpenAI is building the same capability. The article demoes a real encoded ID recovered from a marked PNG via frequency-domain pixel changes.

The thread’s best exchange was about verification. Someone proposed asserting byte-for-byte identity with the last trusted stage of production — a camera you’re sure doesn’t mark, an editor you’re sure doesn’t mark — which sounds right and mostly isn’t, because in the SynthID case the marking happens in the generator, so there is no clean comparator to diff against. You can verify the absence of a specific mark if you know the mark; the point of steganography is that you don’t. A commenter then pushed further than the article: when someone else controls distribution, they control the mark on each copy, so the same innocuous image served to two recipients can carry two different payloads. That converts an origin-tracing scheme into a receiver-tracing scheme, and it’s the version of this that should worry people, because it’s the one that gets built into social platforms. The precedent already exists — Facebook has embedded IPTC metadata tags for years so that images shared off-platform can be traced back.

One pedantic correction the thread got right: spymarking is an application of steganography, not a new name for it. That distinction matters. The authors aren’t claiming to have invented the technique. They’re claiming that “watermark” has been doing rhetorical cover work for it, and they’re right.


Transformers Explained Visually

584 points · Polo Club (Georgia Tech)

An interactive visualization of GPT-2 small — 50,257-token vocabulary, 768-dimensional embeddings, 12 attention heads, 12 blocks — that runs the actual model in the page and lets you step through tokenization, embeddings, multi-head self-attention with live Q/K/V matrices, the softmax and scaling steps, MLP, residual and layer norm, and the final probability distribution. You type a prompt, hit generate, and watch which token wins and by how much. It’s not new (it’s been circulating for a while), but it is the best free artifact of its kind and it earned the front page on merit.

The comment worth repeating, from someone who clearly does this for a living: the most under-explained thing about attention heads is that once the attention matrix is computed, multiplying it by the value vector is exactly a dense layer — the attention matrix is that layer’s weight matrix, constructed at inference time from the queries and keys. So a transformer block is a network that decides its own dense-layer weights, per input, on the fly. That framing rarely appears in beginner material, and the reply explained why: it’s only illuminating to someone who already knows MLPs well but not transformers, which is a shrinking population. You’re either new to both or you know both. Also obligatory and correct: Schmidhuber’s priority claims on this (arXiv:2102.11174), raised in the thread as tradition demands.


Apple has added persistent ‘ads’ to iOS, and it’s driving users crazy

482 points · TechRadar

Promotional banners now sit near the top of the main Settings screen pushing iCloud storage upgrades, Apple Music, AppleCare warranty extensions and TV trials. The reported problem is not that the promos exist — Apple has done that for years — but that many lack a working dismiss control, so users either subscribe or wait weeks for them to expire. The sharpest detail is that paying subscribers are seeing promotions for services they already pay for. The accompanying complaint, from 9to5Mac, is a red badge on the Settings icon for an “Activate iCloud Storage” alert that persists even for accounts with terabytes of free space, navigates to a purchase screen when tapped, and doesn’t clear when dismissed. 9to5Mac’s read is that this particular one is a bug, not a strategy.

I’d take the bug read for the badge and the strategy read for everything else, because Apple’s own public history is unambiguous. Steve Jobs announced iAd in 2010 with the claim that it offered “the interactivity of the web” without hijacking users out of their favorite apps. The program’s failure is now regularly cited as evidence that Apple once had taste; the thread’s sharper version is that Apple did try to make ads tasteful and it didn’t make enough money, so the constraint was never aesthetic discipline — it was revenue. The practical advice that went viral in the thread is the truest thing here: you can bypass the App Store homepage entirely by long-pressing the App Store icon and choosing Updates from the context menu.

The “just switch to Android” subthread was correctly dismissed — Android’s ad situation is worse, and the people with the strongest position were the ones running GrapheneOS or LineageOS with a sandboxed Play Store, which is not a migration path for anyone who isn’t already a power user. The structural complaint underneath all of it is that Apple’s premium positioning and Apple’s ad revenue are in direct conflict, the high-end is where ads are worth the most, and there is no longer a person at the company whose job is to say no.


OpenAI GPT-6 Astra breaks Enigma message that has resisted solution since 2005

453 points · Crypto Cellar Research

On 15 September 2026 Carter Leffer contacted Crypto Cellar to validate his break of German Army Enigma message MVUEH, sent 10 July 1941 by a station with callsign 2ny and received by the SS-Totenkopf Quartiermeister’s radio station, logged as incoming message Nr. 172. It had resisted all attempts since 2005. A sibling message from the same day, Nr. 173 (SIPVX), was broken in 2017 by Alex Shovkoplyas — using the same wheel order (512) as the daily key but different plugboard and ring settings — and that key did not break MVUEH. MVUEH’s key turned out to be entirely different, wheel order 253, and the plaintext is almost identical to SIPVX’s, differing by twelve letters: an enciphering error that turned Bitte into Btte, and the sender repeating his signature (Waschbusch) in the SIPVX message. Two incidental complications: the original ciphertext transcription contained several errors, and the Enigma’s left-hand wheel makes a turnover at the 72nd letter, which is rare and known to make breaks much harder.

The part that matters: Astra did it unprompted. Leffer asked it to see whether it could break any of the unbroken messages published on the site. It analyzed the set, chose MVUEH as the most promising, hypothesized that Nr. 173’s plaintext was related, tried a range of approaches, settled on using the repeated place name ROSENOW ROSENOW as a crib, wrote its own Enigma simulator and Bombe in Python and C++, and ran the search to a correct key and plaintext. Crypto Cellar is still analyzing the logs.

The thread’s best questions were about the crib, which the article did under-explain. Why would a model guess that a place name appeared twice? Answer, from the comments: Rosenow is both a municipality and a district inside it, so a location report of the “New York, New York” form is natural, and the same repetition was present in the 2017-broken message. And the reason a longer crib is worth more than a shorter one: Enigma cannot encipher a letter as itself, so you slide the crib along the ciphertext looking for an offset where no letters align — a four-letter crib has many false positions, a fourteen-letter one has almost none.

I’d resist the framing the headline encourages. The model didn’t out-think Bletchley Park; it found and exploited a structural relation between two messages on the same day, which is exactly the move a human cryptanalyst would make first and which was presumably missed because nobody had considered that the two keys could be that different. The genuinely notable part is the engineering: it wrote the simulator and the Bombe. Models have been able to do that for a while; models choosing the search strategy and then executing it end-to-end on a problem with a verifiable answer is the thing that has changed.


NASA’s Mars Sample Return mission is dead

443 points · Science

Worth flagging up front: this is a January 2026 article that resurfaced. The content is still current because nothing has replaced it. Congress’s compromise FY2026 spending bill backs the White House’s effort to end the Mars Sample Return program, with the joint explanatory statement saying flatly that the agreement “does not support the existing Mars Sample Return (MSR) program.” It does not prohibit sample return: $110 million moves into a new “Mars Future Missions” line covering technologies MSR was developing — radar, spectroscopy, entry/descent/landing, and translation of precursor tech — which leaves the door open for a future Congress to fund a different architecture. NASA’s overall budget landed at $24.4 billion, rejecting most of the administration’s proposed 24% cut, with $7.25 billion for the Science Mission Directorate, a 1% decrease.

The context is a program that got too expensive to defend. Cost estimates rose toward $8–11 billion with return slipping as late as 2040; a 2023 independent review said the architecture had unrealistic cost and schedule assumptions; NASA restructured and narrowed to two landing options at roughly $6–7 billion each and planned to choose in the second half of 2026. Congress removed support before that decision could become a funded mission. Two dozen-plus rock cores sit in Perseverance’s cache in Jezero Crater with no approved retrieval mission, no schedule and no budget — and ESA, which was building the Earth Return Orbiter, has signalled it may rework its spacecraft into a standalone Mars orbital geology mission, which would raise the price of any future revival.

The thread’s argument was the one that always comes up and is always worth re-running. Why not send better robotic analyzers instead of hauling rocks home? Because the analyzers don’t fit. A synchrotron — required for non-destructive nanometer-scale 3D imaging of the samples — is three to four football stadiums in size. Sample return isn’t romanticism; it’s the only way to apply instruments that will never be flight-qualified. There’s also a pointed note from someone who worked on ExoMars: Rosalind Franklin was supposed to launch in 2018, slipped to the early 2020s on a Russian rocket, slipped again, now 2028. Programs don’t die from a single cancellation; they die from deferral accumulating faster than anyone can restart them. Meanwhile China’s Tianwen-3 targets a 2028 launch and 2031 return, at a less scientifically promising site with fewer samples. If it works, they win the race and the race was never the point.


Can gzip be a language model?

359 points · nathan.rs

Yes, sort of, and the demo is more convincing than it has any right to be. The premise is the compression–prediction equivalence from Language Modeling is Compression (arXiv:2309.10668): every prediction model is a compressor and every compressor is a prediction model, because the number of bits needed to encode a symbol is −log₂p and any compressor therefore contains a probability model whether or not anyone wrote one down. DEFLATE compresses the next bytes by finding matches against a 32 KiB sliding window, so a continuation that echoes text already in the window is nearly free. That gives a scoring function: score(candidate) = len(gzip(context + candidate)). Prime the window with a corpus, enumerate candidate continuations, keep the ones that compress best, repeat. Primed on Tiny Shakespeare, the output is recognizably Shakespeare-shaped — correct speaker-tag formatting, period-correct contractions, quoted dialogue — attached to words that don’t exist. It knows something real about the distribution without knowing anything about English.

Vendor-claim skepticism doesn’t apply here because there is no vendor, which is refreshing. The thread’s contribution was mostly historical and useful: the smallest-compressed-size trick is a working classifier (gzip -9 sports.txt test.txt versus politics.txt versus business.txt — the smallest output names the topic), which traces to Witten’s group at Waikato and is the ancestor of the Hutter Prize; and one commenter built a language detector the same way two decades ago by seeding gzip dictionaries with Wikipedia articles per language and picking the best compressor. Both are “not the best approach, but fast and trivially simple,” which is the correct self-assessment. The deadpan best comment: a joke post pitching “quantums” and “Expansive Dictionary Models” that a later reply correctly decoded as a description of BPE tokenization.


AI Has No Wisdom and Neither Will You

349 points · Alex Nedelcu

Nedelcu opens with three things he’s heard in the last month: “I haven’t written code since 2025,” “code reviews are dead,” “people no longer read code.” His case for why that’s a trap rests on a single asymmetry: code maintainability and architecture quality have no reward signal available on any timescale a training run can measure. It takes months or years to observe the cost of bad architecture. Reinforcement learning needs feedback that arrives immediately. So there is no way to train for maintainability — not because it’s subjective, but because the payoff is too far away to attribute. What the model learns instead is the rulebook written for beginners and the patterns present in code in the wild, and most code in the wild is bad. His inline examples are concrete: models that “simplify” by extracting functions that are neither reusable nor clarifying, so understanding the outer function now requires reading the inner one — which is a net increase in work disguised as a refactor.

The line worth remembering is his definition of expert intuition: it’s built from long hours debugging production, it’s context-dependent, and it resists being written down as rules. “Experts don’t follow the rules, they make the rules.” Which means the artifact a model trains on — the rules — is by construction the beginner’s layer, and the layer above it is exactly the part that doesn’t serialize.

The thread’s best thread was the manufacturing analogy, and it’s a genuinely good one. When the West offsourced manufacturing, the dismissals were confident and theory-backed: comparative advantage, Ricardian tables, everyone-wins. The people making the argument weren’t lying and the theory wasn’t stupid; the institutional capability eroded anyway, and the erosion was invisible while it happened because each individual outsourcing decision was locally correct. Two commenters described being young and ideologically on the opposite side of that argument and being unable to articulate why it was wrong, only that their lived experience said it was. Applying that to code: the individual decision to stop reading code is locally correct, the aggregate consequence takes years to become visible, and by then the people with the intuition to run the review have left. Also noted, correctly, by an AI-coding skeptic in the other model threads: AI is good at burning down accumulated tech debt that nobody had bandwidth for, which is a real and underrated win and doesn’t contradict any of the above.


Turning Apple Intelligence off (and getting 22 GB back)

339 points · Apple Support · 235 points · r/MacOSBeta

Two stories that belong together, because the second is what the first implies. Apple’s own support document is the canonical list of off switches on macOS 27, and reading it is the point: Siri AI (Beta) off via System Settings → Siri → Turn Off Siri; Siri Classic as a substitute; Messages summaries in Messages → Settings; email summaries under Mail → Viewing; notification summaries under Notification settings; personalized Smart Replies under Mail → Composing; voicemail suggestions under Phone → Calls; Journal writing prompts; and for anything left, Screen Time restrictions. That is ten separate surfaces, several of them inside individual apps, and the catch-all is a parental-controls feature. Apple’s own documentation calls the section “turn off and restrict,” which is an honest description of the design: “restrict” means the model still ships, still occupies the disk, and is merely hidden.

Which is where the Reddit thread comes in. The workaround people are sharing exists because there’s no supported way to remove the models — only to hide their UI. Commenters report recovering 16+ GB by deleting the model directories, and one proposed the more robust version: after deletion, create a zero-byte file with the same name and set the immutable flag (chflags schg) so the redownload fails silently instead of succeeding on the next check-in. Another thread asked the obvious structural question and got no good answer: macOS 27 “Golden Gate” landed as a ~28 GB update, adds a liquid-glass redesign, and users report fewer working programs, worse contrast and readability, larger but less distinct tap targets, and no feature they can name. The AI models are the only part of that update with a measurable footprint, and the only part you cannot delete through the interface.


AI coding has made CI a bottleneck, so we reworked ours to keep up

305 points · Linear

A concrete engineering post from Linear’s Mufeez Amjad, framed by the observation that agents made shipping faster but validating didn’t, so CI became the constraint. Four workstreams: move workloads off GitHub Actions onto third-party runners with faster CPUs, better storage and better cache; optimize the jobs that gate everything else; reduce repeated setup; make test execution itself more efficient. Results: test suites nearly quadrupled since January, and PR wait time went from over six minutes to just over five, with runner time per test roughly halved. They’re candid about the shape of the graph — machine time per test spikes whenever they added shards, which shortens the wait at the cost of more machine time, so it’s a visible tradeoff rather than free improvement.

The thread’s objection is the one worth taking seriously, and it wasn’t answered: if every team is moving this fast, why aren’t the products better? The specific complaint was that the new iPhone and Android ship with fewer features than usual, no indie team has shipped a Linux-sized alternative OS, and Windows takes three seconds to open the right-click menu. A reply from inside a company doing this work gave the honest accounting: a big chunk of AI effort went to quality and availability improvements and to tech debt nobody had bandwidth for, QA is catching issues earlier, and CI went from roughly 1.5 hours to under five minutes — and internal costs are rising directly from AI spend, to the point where they expect it to become a new per-employee line item. That’s a real answer to “where did the gains go,” and it’s an uncomfortable one: the gains landed in internal process and reliability, not in things users can see, and they arrived with a new recurring bill. The counter-reply — if users don’t perceive the improvement, it’s legitimately “sameish” — is also right, and the two positions aren’t reconcilable because they’re measuring different things.


Heretic removes restrictions from language models

263 points · heretic-project.org

A polished front-end for abliteration — the technique of removing refusal directions from a model’s weights rather than prompting around them. The pitch is explicitly aimed at non-specialists: no ML PhD, no expensive hardware, pip install -U heretic-llm then heretic Qwen/Qwen3.5-4B, and it’s fully automatic from there. Under the hood it identifies “abliterable components” (o_proj and down_proj in the tutorial’s model, 32 modules each), loads 400 good prompts from mlabonne/harmless_alpaca and 400 bad ones from mlabonne/harmful_behaviors, computes the refusal direction from the difference, and writes LoRA adapters. The VRAM rule of thumb is ~2.5 GB per billion parameters, with bitsandbytes 4-bit quantization cutting that by ~70% — an RTX 3060 handles a 4B model. It’s built on prior work (the diff-in-means approach), but the packaging is the contribution.

Vendor-claim skepticism is beside the point; the thread argued about whether the tool should exist. The strongest case for it is a commenter with a Chinese IP camera that has known CVEs and no vendor support, who says no mainstream model will help with the reverse engineering needed to reclaim control of hardware they own, so an abliterated model was the difference between using the camera and shipping it back. The counter-argument — that safeguards exist to raise the cost of attacks — is met by a claim several people made independently and which I find persuasive: Kimi K3, GLM-5.3 and Qwen3.8-Max are near-frontier and will happily do security work and reverse engineering, so the safeguard your main model enforces is not a control on adversary capability, it is a tax on legitimate defenders. One caveat on that: “I asked GLM and it agreed” is an anecdote about willingness, not competence, and the two get conflated constantly in these threads.


Python Workers are now generally available

261 points · Cloudflare

Two years after the preview, Python on Workers is GA — meaning first-class, fully supported, able to reach Workers AI, R2, D1, Hyperdrive, Durable Objects, Queues and Workflows, and able to run FastAPI, Django and Flask via an ASGI entrypoint, including nested Dynamic Workers. The runtime is a Wasm-compiled CPython (Pyodide) running on the Worker isolate model, which is what makes it possible and also what makes it not-quite-Python.

The thread produced the most technically substantive exchange of the day, and it’s about maintenance funding rather than capability. A urllib3 maintainer explained the provenance of the Pyodide support that makes the Requests library work: large contributions were merged, along with JSPI support, and the funding for that work went to the external contributor who implemented it rather than to the maintainers. Now the project owns code it didn’t write, gets the support burden, and — the maintainer’s framing — there’s a meaningful difference between funding a contribution upstream and funding the people who have to maintain it afterward. The suggested resolution, funding upstream maintenance directly with an internal champion, is the correct answer and rarely happens. Second issue, from the same commenter: Pyodide and Cloudflare don’t use the real network stack, patching HTTP to go through JS fetch and patching Python’s event loop to sit on the JS event loop. A Cloudflare engineer replied that the term for the JS model is “eager,” not “preemptive,” that Python coroutines stay lazy under the WebLoop, and that Pyodide can use direct sockets in Node and Cloudflare — but not in browsers, where the same-origin model forbids it. Which is a fair correction, and leaves the portability complaint standing: WASM Python is semantically close to CPython and not identical, and every difference is a bug you will find in production.


I asked Meta’s Muse for its filesystem and it sent me 6.8GB

260 points · mouse.dev

The author asked Meta’s Muse agent (internal name Hatch) to archive the files it could see and deliver them to Google Drive. It did: ~2.7 GB compressed, 6.8 GB unpacked, containing what appears to be the root filesystem of the Linux environment his session was assigned — Ubuntu system files, Muse’s internal documentation, integration code, app templates, memory files, agent logs, and SSH key files. The interesting paths are /home/hatch (containing SOUL.md, IDENTITY.md, USER.md, MEMORY.md, AGENTS.md, TOOLS.md, plus agents/, docs/, memory/bank/, dreams/, workspace/self_improvement/) and /opt/hatch/{skills,runtime-cell} and /opt/hatch-image/bin/codex with a bwrap sandbox binary. Roughly twenty Markdown files described browser use, connectors, payments, credentials, data handling, generated files, voice, goals and scheduling.

Being careful about what’s actually established, since the post itself is careful: the author reported this through Meta’s bug bounty program, did not publish the archive, the keys, or the session logs, and explicitly did not demonstrate a container escape — a chat message in one of the screenshots claims one, and he flags it as an unverified claim by the agent. He also hasn’t established whether the SSH keys were live or what they’d have reached. The reported concern is narrower and sharper than “container escape”: sensitive internal runtime material was exfiltrable through an ordinary conversation plus a connected third-party export destination, which is a capability most agent products have and most threat models don’t cover.

The thread’s best comment is a two-parter that gets at something real. First: “the markdown is the product, not the process” — a large fraction of a modern agent product is a directory of natural-language instructions, hand-tuned or model-generated, and there’s no way for a customer to distinguish a rigorously evaluated system from an accreted pile. Second, and the reply that complicates it: unreadable skills don’t imply proper engineering either; they’re just as consistent with a blind iterative loop with no improvement signal. Both are true, and the reason the story is on the front page is that a filesystem dump is currently the only way to tell which one you’re getting.


M5 Ultra Mac Studio Review

260 points · MacStories

A review of the 256 GB M5 Ultra Mac Studio, framed as “the dream Mac for local AI agents” — the argument being that a machine with enough unified memory and reasonable bandwidth changes what a local model can do, so skepticism about pairing an agent framework with local models becomes outdated. The comparison set is an M3 Ultra with 512 GB and a desktop with an RTX 5090. The reviewer’s conclusion is that the 5090 still wins on bandwidth but that for size, thermals, noise and being a Mac, he’d take the Studio.

The numbers people actually came for are in a chart near the bottom, and a commenter extracted them. Qwen3.8 27B generation speed in tokens/sec at 8K / 64K / 128K / 256K prompt: RTX 5090 PC 59 / 51 / 44 / n/a; M5 Ultra 48 / 39 / 32 / 24; M3 Ultra 31 / 23.5 / 20 / 15. The M5 Ultra’s story is the long-context column — it keeps generating at 256K where the 5090 has no entry, which is the real argument for unified memory over a discrete card. But the 5090’s absolute numbers drew immediate objection from people running them, and this is where the review’s framing is weakest: multiple commenters report the 5090 sustaining 200+ tokens/sec on the same model with NVFP4 quantization and multi-token prediction, which is one to two orders of magnitude off the chart. The plausible explanation is a stock llama.cpp-class runtime versus an optimized one, and the reviewer didn’t say which stack he measured — for a review whose headline claim is about local AI capability, the inference stack is not a detail. A good rejoinder in the thread, though: raw throughput isn’t the whole question, and one 5090 owner described the card as “the worst ADHD team member” — fast, great for shallow precision work, but needing constant supervision from a larger model, which puts the time-to-result much closer than the tokens/sec ratio suggests.


Raspberry Pi blocks changing RAM chips

247 points · Raspberry Pi forums

A Raspberry Pi engineer’s forum reply, now widely circulated: the foundation has been “locking devices to their original RAM size” in firmware for some time, and swapping RAM chips between boards won’t work, including in some cases same-size swaps, because chips are programmed with the timing and capacity settings of the specific part they shipped with. A failed check produces error code 9 and a refusal to boot. The rationale, from the engineer: buy a 2 GB board, swap in a cheap 8 GB part of dubious provenance, resell as an 8 GB device — the board hasn’t been validated by Raspberry Pi, the failure modes are weird and intermittent, and the support burden lands on them. Removing the commercial incentive removes the fraud.

The two facts that change how you read this both come from the community, and both are in the coverage. First, the firmware change landed between 10 and 23 September 2024 — a year before the current RAM shortage, so the “they’re reacting to RAM prices” theory is wrong, though the incentive to counterfeit boards is obviously much higher now. Second, it’s bypassable: the bootloader lives in SPI flash and can be reflashed over USB with rpiboot, and nothing stops you flashing the pre-September-2024 version, so anyone who can perform a BGA RAM swap can perform a bootloader downgrade. Hackaday’s position — that this barely affects repair, that the main audience for RAM swaps turned out to be dodgy resellers, and that 15 separate people on GitHub complaining about bricked boards they bought on AliExpress or Amazon is evidence the fraud is real — is the strongest contribution to the thread. Jeff Geerling’s counter-proposal is also right in principle: older Pis shipped a warranty bit that flipped when you bypassed safety limits. Set a one-time flag declaring the board out of warranty, let the owner do the mod, and put them on the hook. That gives up the enforcement entirely, which is exactly why it wasn’t chosen — but it also gives up the fraud detection, which is the only thing the current design does well.


AMD’s random number generator can’t generate a 0?

233 points · flat assembler forum

The premise sounds like a joke and isn’t. A user rendering random data into charts via rdrand and rdseed noticed that on AMD hardware the zero bucket in a 16-bit space is always empty, while the same program on Intel produces zeros fine. Reproduction across the thread narrowed it precisely: it’s rdrand16/rdseed16, the 16-bit operand forms; 32-bit and 64-bit forms are unaffected. And it’s not that zero can’t be produced — it’s that when a zero is produced, the instruction sets CF=0, which is defined to mean failure, retry. Since every sane RNG wrapper retries on CF=0, and the retried result is again a legitimate value from a 1..65535 distribution, programs that request 16-bit random values never see a zero. On Intel, the same code sees zeros at the expected 1/65536 rate.

The erratum is public and the history is worse than the bug. AMD-SB-7055 documents RDSEED failure on Zen 5: rdrand16/32 return zero with CF=1 on entropy exhaustion, and AMD’s recommended handling — treat an all-zero result as failure and re-roll — recreates the exact behaviour being complained about, because it converts a legitimate zero into a retry. A commenter points out this is plausibly how the Zen 1/2 problem was created in the first place, and AMD says it may be addressable by a future microcode update. The measured statistics from one reproducer are worth quoting because they close the case: 1 billion rounds, 15,312 CPU-flag failures, and zero cases where the failure carried a non-zero result — every failure was a discarded zero. Bucket counts across 1..65535 are uniform to within noise, the zero bucket has exactly the failure count in it, and there is no way to recover the value. This is the small, boring, almost-unnoticeable class of bug that matters more than it looks: an instruction that lies about its output status, a documented handling recommendation that institutionalizes the lie, and a fix that requires a microcode update most affected machines will never receive. Also, the thread’s earlier report is worth noting — Zen 2 once returned all-ones from rdrand until a microcode fix, so the family has form here.


US halts flights at busy East Coast airports

231 points · Reuters

On Monday 21 September the FAA issued ground stops at five major Northeast airports — JFK, LaGuardia, Newark, Teterboro and Philadelphia — plus Boston Logan and for a period Washington National, all served by the Philadelphia TRACON facility. Delays averaged 44 minutes at LaGuardia, 105 at Boston and 212 at Philadelphia. The sequence, per Transportation Secretary Sean Duffy and FAA Administrator Bryan Bedford: an aging circuit at Philadelphia TRACON, scheduled for replacement, failed. The facility moved to fail over to its backup and discovered the backup fibre had a break — a fibre cable running adjacent to an Amtrak rail line in New Jersey, cut by a construction crew working in the area, roughly 600 feet of it somewhere between New Brunswick and Newark. Verizon, which owns the cable, stated its facilities were fully functional until the cut and that it was not responsible for the damage. Estimated repair time: about 13 hours. Ground stops were lifted and operations resumed the same evening.

The detail that makes this more than a travel story is that it coincided with a GAO report finding the FAA has not completed risk assessments or updated its security documentation covering spoofing, jamming and other radio-spectrum attacks on the systems guiding US aircraft — which prompted a senator to describe it as a major threat to national security. To be precise about what happened: this was a physical failure, not an attack, and conflating the two would be sloppy. But the outage is a live demonstration of the property the GAO is complaining about — a single aging circuit with a backup that had already failed, no monitoring that surfaced the dead backup before it was needed, and a dependency on a third party’s fibre trenched alongside a railway that anyone with an excavator can reach. Redundancy that exists on paper and hasn’t been tested in practice is not redundancy. Two independent failures of the same dependence is the definition of a single point of failure that nobody had modelled.


Apple iPhone 18 Pro Camera test

205 points · DXOMARK

Overall score 172. The hardware is three 48 MP sensors: primary 24 mm f/1.48–f/4.0 with variable aperture, sensor-shift OIS and Focus Pixels; ultra-wide 13 mm f/2.2 at 120°; telephoto 100 mm f/2.8 in a tetraprism design with 8x optical zoom and sensor-shift OIS. DXOMARK’s summary credits wide dynamic range, improved contrast and skin tones, effective stabilization, and a well-balanced texture-noise tradeoff, and calls the variable aperture a particular strength because it maintains sharpness in complex scenes with multiple subjects — which is the real point of a variable iris in a phone: it’s about retaining focus across a scene, not about depth of field control. Flare is better controlled than the previous generation with reduced diffuse flare, though green spots still appear and flare is visible when the iris is closed.

Two caveats, because the report invites them. First, DXOMARK says explicitly this is a short version and the full review is coming — so the 172 is not final in the sense that the reasoning behind it hasn’t been published. Second, an overall score of 172 is an aggregate over photo, zoom and video sub-scores, and the sub-scores are where the information is; a single number with no comparison set on the page is a press artifact, however rigorous the lab work behind it. The genuinely interesting item is the variable aperture, because it’s the only component here that changes the physics rather than the processing, and it’s the one thing the summary attributes a concrete improvement to.


OpenAI is well positioned to fast-follow Jev

204 points · arcturus-labs.com

The argument is narrow and mostly correct. Jev — TypeSafe’s decision model, reported by Vercel as the fastest-adopted model in AI Gateway history — is, per the author’s assumption, a conventional LLM used as a classifier: run one forward step, take the probability distribution over the next token, and read off the candidates you care about. For a yes/no question, look at the probabilities of true and false and normalize. For a multiple-choice question, look at the tokens for the option names. That’s the whole trick, and the evidence for the assumption is that many of the early clones are themselves LLM-based. If that’s right, OpenAI needs to replicate the training, not invent a capability, and OpenAI has spent years using LLMs as implicit classifiers — for routing, for safeguards, for data preparation — without ever packaging general classification as a product. The author’s thesis is that replicating it would let OpenAI not just clone Jev but embed the classifier inside existing models and agents for model selection, cheaper thinking, guardrails and latency, which TypeSafe has no path to.

Where the thread pushed back, and it’s the right pushback: the framing understates how much of the market already exists. Every large lab has in-house classifiers of every size and specialty; the question is which ones are worth exposing as an API, and “none, mostly” is a business answer rather than a technical one. Someone also noted that OpenAI did ship a general-purpose zero-shot classifier API on GPT-3 around 2020–21 and nobody cared, which is the strongest counterexample in the thread — the capability existed, the market didn’t. The most useful contribution came from people who tested the models: Jev and Laya are competent but a task-specific classifier beats them, and getting that classifier is no longer hard — upload a spreadsheet, get a scikit-learn model that’s free, instant and more accurate. That’s the same objection the Kev thread made yesterday and it hasn’t been answered: the expensive part of building a classifier used to be expertise, and expertise got automated, so the remaining differentiator has to be single-call heterogeneity plus calibrated uncertainty, and nobody has demonstrated that’s worth a checkpoint. Fair-minded caveat on both directions: the skeptics are people who tried the free tier, and the boosters are people who shipped something.


Divide by depth for instant 3D

203 points · gabrieloc.com

A short, well-made explainer on the projection at the core of every 3D renderer: given coordinates where y is up and z is forward, project to 2D with x' = x/z and y' = y/z. Points farther away converge toward the vanishing point at the origin; the table in the post shows (2,1,2)→(1,0.5), (2,1,4)→(0.5,0.25), (2,1,8)→(0.25,0.125). The demo is a ball orbiting the camera’s up axis, offset along z by a constant, projected with the same two divisions and scaled by radius = 0.5 / z, followed by a wireframe cube built from twelve edges each projected endpoint-by-endpoint, in about twenty lines of shader code. The point of the post is the one the author makes about his own learning: cameras in high-level frameworks “just work,” which means you never learn the vocabulary to search for what you want until you write the low-level version, and the low-level version turns out to be two divisions.

The thread’s correction is worth including because the post’s framing invites the error. Several commenters hammered the distinction between linear and affine: translation is not a linear transformation in 3D because it doesn’t fix the origin, so it cannot be represented by a 3×3 matrix; embedding 3D in 4D at w = 1 is what makes translation representable as a matrix, which is what makes a single 4×4 composition possible. The nuance a second commenter added is the one that actually matters for performance: you do not need homogeneous coordinates to compose a series of linear transforms — composing linear maps is just linear algebra — you need them only to include translation in that composition. And the parenthetical historical note is the best line in the thread: the win was largest on ancient hardware where every MUL was contested, and the modern machine mostly doesn’t care. Then it wandered to quaternions and gimbal lock, which is where every graphics thread ends up, with a genuinely good anecdote attached — Collins asking for a fourth gimbal for Christmas, four hours before Armstrong’s step.


Still on the page

Thirty-one of today’s fifty-six stories above 200 points were covered here in the last four days, and the frozen-front-page pattern is now the most reliable thing about this feature. The largest single mover was a piece from three days ago: AI-generated posters don’t have to be horrible 1,849 → 1,870, still the top item on the page. Attention is all you have 386 → 1,016, up 630, which is what happens when an essay about feeds gets fed to a feed.

Other notable climbs: I built non-autoregressive decision models with RL a year ago 1,324 → 1,333 — a priority dispute still accumulating points. Android 17 without an AOSP release 1,170 → 1,177. What happened to the Snowden archive 653 → 710. What Sun got wrong 394 → 666, up 272. Grok 4.7 363 → 596. Fable 5, median thinking declined in August 251 → 414. Kev 343 → 449. ZuckOff 340 → 397. Grim Fandango Puzzle Document 343 → 372. I am often wrong 321 → 333. MCP was always a bad idea? 312 → 329. The Effect of CRTs on Pixel Art 301 → 308. Singapore’s reading micropayments 279 → 286. Why do we need human mathematicians anymore? 275 → 288. The senior engineer death spiral 238 → 260. Mini-AGI 236 → 270. The LLMentalist Effect 220 → 230.

Also still up: Qwen Image 2.1 731, Exfiltrate your Weights 730, How to Write with an LLM 751, Two parallel neural ectoderm progenitors 659, AX 652, Samsung HBM4 554, Spain blocks Archive.today 544, If math is more than proof 429, What Zig felt like 275, RSA-896 226, Measure internet censorship 221.

Two stories that were covered here dropped out of the 200+ set entirely and are worth noting for the deltas: Disney+ ad policy 474 → 504 and Jev-Leftpad 226 → 230 — both still gaining, both now ranked below the window this scan covers.

Throughline

The cost curve is the product, and the disclosure page is the marketing. Three model releases in one day, and not one of them leads with capability. Anthropic’s Opus 5.5 announcement opens by reminding you of last week’s call to pace the frontier and then spends the rest of the page cutting prices — 40% cheaper to run, cache reads down 60%. OpenAI halved Sol and Luna against its own promotional pricing and framed it as “distributing the benefits” of Astra’s intelligence. Xiaomi’s MiMo leads with a composite index score but the thing that actually impressed people was a live dashboard of the training run. Set beside yesterday’s Fable 5 measurement — the same model name serving materially different inference regimes — the pattern is that vendors have stopped competing on what the model can do, because that’s become unverifiable at the point of sale, and started competing on what it costs per completed task, which is verifiable. The MiMo dashboard is the only item in the set that lets a customer check anything, and that is precisely why it’s the most-discussed release of the day despite having the lowest profile.

The off switch is being removed, one surface at a time, and none of it is announced. Apple deleted the Apple Intelligence toggle and hid the rest under Screen Time, on a machine where the models consume 22.28 GB and cannot be removed through the UI — and separately shipped promotional banners in Settings with no dismiss button, including to people already paying for the service being promoted. Google’s SynthID, as characterized by the spymarks essay, embeds a 64-bit database identifier that maps to a user record into generated media by design. Raspberry Pi’s bootloader refuses to boot a board with a RAM chip it didn’t ship. AMD’s RNG silently reports a legitimate zero as a failure so that retry loops never see one. Meta’s Muse will archive its own filesystem and hand it to a third-party destination if you ask nicely. These look like six unrelated stories and they’re one story: the default is yes, the refusal path is either hidden, undocumented, or absent, and the failure mode is always a mislabeled status bit — the AI toggle that says off, the RNG that says error, the RAM check that says corrupt, the badge that says dismissible. None of these are malicious in isolation. All of them require the user to detect a lie about state.

The measurable work is the work that gets automated, and the work that isn’t measurable is the work that matters. Colin Breck’s essay and Alex Nedelcu’s are making the same argument from opposite ends: AI is superb at verifying a paragraph against a log, completing a citation, writing an abstract, and burning down accumulated tech debt nobody had bandwidth for — and it cannot write the paragraph that transfers the context only you have, because that context isn’t in the prompt, and it cannot learn maintainability, because the reward arrives years after the run ends. Linear’s CI post is the honest bookkeeping: test suites quadrupled, wait times came down, runner time halved, user-visible features did not improve, and internal costs went up enough to become a new per-employee line item. And the day’s best counterexample to the pessimism is the Enigma break, where the model did something nobody had done since 2005 — by first writing itself an Enigma simulator and a Bombe. Astra did the boring, measurable engineering before attempting the unmeasurable part, which is the only order that has ever worked.