Forty-one stories cleared 200 points across the rank window and a 40-hour date sweep. Twenty-six of them are new. The rest are the same items this roundup has tracked since the weekend, and several of them moved a long way — the Claude Haiku launch nearly tripled its score overnight.

The page splits cleanly in two today. Above the fold it is a memorial. Below it is the most interesting twenty-four hours the AI-maths story has had yet: a release, a critique from the most respected mathematician alive, a retraction, and a paper arguing that the verification everyone is pointing at does not actually verify what they think it does.

Margaret Hamilton has died — 2,058 points

2,058 points · 247 comments · news.mit.edu · HN discussion

Hamilton died at 90. She led the MIT Instrumentation Laboratory teams that wrote the flight software for both the Apollo Lunar Module and the Command and Service Module — more than 400 people — and the photograph most people know her from is her standing beside printouts of that code, stacked roughly to her own height.

The technical legacy is specific and worth stating plainly, because “led the software team” undersells it. Hamilton coined the term “software engineering” to describe work that was, at the time, treated as an appendage of hardware. She and her team built priority-based scheduling and error detection into the Apollo Guidance Computer rather than assuming a perfect run, which is why the 1201 and 1202 program alarms during Eagle’s descent — the computer shedding low-priority tasks under load — surfaced as a recognisable alarm instead of a hang. Mission Control had a procedure to interpret them. The landing continued.

The thread is mostly people who met her. One commenter describes a VC firm founded by Apollo-era Draper Lab people, and a conversation with Hamilton about formalised control systems that he understood none of. Another points at the Computer History Museum’s oral history and connects a story in Levy’s Hackers about a late-night TX-0 session that corrupted a weather simulation — possibly the code of Edward Lorenz, whose rounding-truncation bug later became the butterfly effect. The undercurrent in the thread is a complaint about how badly most institutions use competent people, which is the right thing to be arguing about on the day the archetype for that problem dies.

Trump administration is suspending Microsoft from a green card program — 690 points

690 points · 1,168 comments · apnews.com · HN discussion

The administration is suspending Microsoft from the permanent labour certification programme — the step where an employer must demonstrate that no qualified American worker was available before sponsoring an H-1B holder for a green card. Vice President Vance announced it, alleging the company gamed the market test. CNBC reports Adobe was suspended at the same time, from the Department of Labor’s PERM programme, with Labor Secretary Keith Sonderling named as the source.

What is actually happening is narrower than the headline. This is an enforcement action against named employers inside an existing programme, not a restructuring of it. The allegation — that companies post a requirement in a small-town paper, receive no qualified response, and treat that as the market test — is the mechanism PERM has always invited, and it is a procedural claim rather than a finding of immigration fraud. The 1,168 comments make it the day’s most-discussed item by a factor of two, and the top comment is Vance’s own description of the newspaper-advertisement mechanism, which tells you the thread accepted that framing immediately.

The substantive argument in the thread is one the immigration debate rarely gets to. An employer sponsoring your green card ties your residency to a single company for years, which is exactly the arrangement people describe as indentured and exactly the arrangement that makes underpayment possible — the top commenter’s point that Microsoft “brought in foreigners and paid them less to do the same job” is a claim about leverage, not nationality. Two commenters posted sponsor-transparency databases (h1bgrader, permtrack) letting you compare per-employer filed wages, which is the most useful thing on the page. The counterargument, almost never made because it is unglamorous: a lottery with no country cap and no employer tie would be both simpler and fairer than anything either administration has proposed, and neither has proposed it.

Show HN: Bigwords.page – Turn any screen into a sign. The URL is the app — 677 points

677 points · 155 comments · bigwords.page · HN discussion

A single-page tool where the message, styling and behaviour all live in the URL fragment. #Hello World renders a sign; &bg=, &fg=, &font= and &interval= configure it; %0A breaks lines; || separates slides; {} placeholders drive countdown timers. Text auto-fits any screen from a phone to a stadium board, QR codes for links and Wi-Fi are generated in the browser, and there is no account or backend.

The privacy claim is the interesting part and it happens to be structurally true rather than marketing: URL fragments are never sent to the server, so the message genuinely cannot reach whoever is hosting the page. That is also the whole business model — there is nothing to host, nothing to store, and nothing to charge for. The author’s own motivating case was a remotely managed tablet where he needed to change a wall display without physical access and without configuring a kiosk browser, which is a real problem with no good existing answer.

The thread’s sharpest note is not about the tool. Someone who built the same thing reports a Firefox scrollWidth bug where text fits on screen but reports overflow, and links the two-line fix — which the author, unable to deploy anything, cannot apply to anyone’s existing links. Someone else flags the Wake Lock API for the keep-the-screen-on problem the site does not solve. And one commenter takes a swing at the tagline: “the URL is the app” is exactly the kind of line an LLM would drop into every other paragraph. Fair. It is also, unlike most such lines, literally true here.

Tell HN: I’ve been paying for a rural Tanzanian’s education for 10 years — 591 points

591 points · 162 comments · news.ycombinator.com · HN discussion

The submitter was 19, spent a summer in Ibumila in the Njombe highlands, and offered to cover the ~$200-a-year boarding school fees of the 11-year-old daughter of the family he stayed with. Anwarite has now graduated from the University of Dar es Salaam in Language Studies and been accepted into an MA in Swahili. Her full 18-month programme — tuition, registration, fees, books, supplies — is 3,060,000 TSh, about $1,161. Roughly 2% of people from Njombe finish a degree. He registered the Tanzania Education Project as a 501(c)(3) in 2021 and now supports five students from the same village.

Read the numbers rather than the sentiment, because the numbers are the argument: a master’s degree in Tanzania costs less than two weeks of a decent US salary, and the binding constraint is the absence of a village school rather than the cost of tuition. Ten years of consistent small transfers beat any one-off large donation, and the mechanism — one person who stayed in contact — is the part that charities with 40% overhead cannot replicate.

The comment thread is the honest version of this genre. The first serious reply asks for the EIN to be posted prominently, because the commenter donates through a Schwab donor-advised fund and could not find the organisation by name; another warns about the long history of child-sponsorship scams and asks for a trusted party to vouch; a third asks whether structural aid (workshops, seminars, travel for trainers) works better than paying school fees directly. The author shows up to fix a typo in the email address on his own site. All of which is to say the crowd’s instinct was verification first, and it was right to be.

“Math 2.0” will need to value mathematical progress more holistically — 575 points

575 points · 597 comments · mathstodon.xyz · HN discussion

Terence Tao’s four-post thread is the single most important thing written about the OpenAI mathematics release, and it is important precisely because it is not a capabilities argument. His claim is about incentives: “Math 1.0” placed a premium on being first to solve an open problem even when the solution was not initially understood, and that objective has now been “optimized to the point of unsustainability.” Math 2.0 has to decentre raw problem-solving and start valuing exposition, community building, and the opening of new directions of study — and the community has to re-evaluate its criteria for education, publication and career advancement to match.

Read that as a commons problem rather than a complaint. The field’s entire prestige economy — who gets the chair, the grant, the graduate students — was priced against solving, and solving just became cheap and instantaneous. Tao’s second point is the one the release page does not address: an AI agent pointed at a list of open problems produces answers and then walks away, with no stake in the field and no interest in the follow-up work that makes a result useful. He explicitly says AI can contribute in the directions he names, but that it “will require more imagination and ambition than the ‘Math 1.0’ mindset of simply pointing one’s favorite AI agent at some set of open problems.”

Two comments sharpen it beyond Tao. The first is the contamination argument: the mere knowledge that a solution exists changes how everyone else approaches the problem, because the search for an alternate route that reveals why the thing is true is no longer the same search. The second is that the mechanism producing these results looks less like reasoning and more like cross-pollination — the model sees thousands of subfields at once and connects results that specialists in either field would never have read. That would explain both the successes and the reported illegibility. If that is the actual capability, “Math 2.0” is not a philosophical question. It is a question about what happens to a discipline whose value was specialised depth at the exact moment depth stopped being scarce.

Animated ASCII Art for Web Pages — 444 points

444 points · 74 comments · ascii.rest · HN discussion

Roughly 190-196 animated ASCII pieces shipped as small typed TypeScript modules with no dependencies, for React, Next.js, Astro or a plain HTML tag, under the MIT licence. The set covers full-colour scenes, loaders, charts, typography, 27 language logos, 22 Linux distributions in either colour or a single ink, plus 3D shapes, physics simulations and creatures.

The design decision worth naming is that these are code rather than images, which means they repaint at the target’s frame rate and resolution instead of being scaled bitmaps, they cost nothing to change, and each is small enough to inline. The single-ink variants for the distro logos are the detail that shows someone thought about actual use — a monochrome terminal-flavoured brand mark is a real need and almost nobody ships one. The counterpoint is that ASCII art is a novelty aesthetic, and novelty aesthetics do not survive contact with a design review.

The Mathocalypse — 367 points

367 points · 383 comments · scottaaronson.blog · HN discussion

Aaronson’s post opens with his nine-year-old taunting his wife, complexity theorist Dana Moshkovitz, with the news that a robot solved the problem she has worked on for her whole career, and then argues that the brat was right: this was one of the biggest days in mathematical history. Among the 372 results released were a proof of Subhash Khot’s Unique Games Conjecture — published on the recommendation of an advisory group including Timothy Gowers and Edward Witten — which underpins a large fraction of known inapproximability results.

Aaronson’s own framing is the generous one, and the comment thread supplies the corrective in three parts. First, the actual hit rate: the model was tried on roughly 8,000 problems and resolved about 5%, which is simultaneously extraordinary for a single three-hour attempt per problem and the honest denominator everybody quoting “372 results” is omitting. Second, the read quality. Aaronson quotes people attempting to digest the proofs, and one says it “feels like something written by someone who’s on psychedelics” — unclear, nonsensical in places, name-dropping prior work without explaining how it escapes known impossibility results. The third is the labour problem and it is the one that connects to Tao: someone has to read these, and the reading is thankless, barely credited, competitive and unfun. A commenter points out this is already the day job of a large number of software engineers — you spend your time reading and verifying code you did not write, and this is that, only with the prestige of mathematics attached.

The best line in the thread is the observation that the post itself is proof of human authorship, because only a human nine-year-old talks like that. The second-best is a pointer to Ted Chiang’s 2000 story The Evolution of Human Science, which is about a world where human comprehension of science has been permanently left behind by its tools, and which has been linked on Hacker News roughly once a month for five years without ever losing its edge.

Meta and Microsoft take steps to reduce employee usage of Claude AI — 364 points

364 points · 379 comments · rswebsols.com · HN discussion

Per The Information, both companies are cutting back internal use of Anthropic’s Claude and shifting toward their own coding tools. The figure that matters is the budget: inside Microsoft’s cloud and AI division, monthly per-employee AI spending limits reportedly went from $100,000 to about $10,000 in most cases.

Treat the sourcing as reported rather than confirmed — this is an SEO-oriented aggregator relaying The Information, and no company has put a number on the record. But if the direction is right, it is the most consequential revenue story of the quarter. A $100,000 monthly allowance per employee was the actual policy at a hyperscaler during the token-maxing era, which means the order-of-magnitude correction is a tenfold cut and the era it ends was real. Add the thread’s second data point — an ex-Meta employee saying Anthropic’s revenue is concentrated enough that two clients could account for a quarter of it — and this stops looking like a procurement decision.

The comment thread converges on a reading that is more interesting than either “Claude got worse” or “Claude got too expensive”: big labs dogfood their own models and block competitors for data-governance reasons, so the switch is partly structural. The other half is the accountants. One commenter describes losing Claude access because of cost and predicts a reckoning over “token maxing”; another reports the emerging enterprise pattern of tight budgets plus an internal API marketplace routing traffic to whichever model is cheapest, including open-weight and Chinese models. Both can be true at once. The frontier labs are being asked to justify per-seat costs that their own sales decks made easy to approve two years ago, and the models are now good enough that the answer to “do we need the best one” is often no.

Living off-grid: Hundred Rabbits — 355 points

355 points · 166 comments · 100r.ca · HN discussion

Devine Lu Linvega and Rekka Bellum have been living and working off-grid for years — currently on a sailboat — and the homepage is a running log of what they have shipped. Recent items include a definitive ten-year re-release of Donsol, their card game, now with animated sprites and a proper manual, runnable in a browser; a second “Rabbit Waves” instalment on encrypting and decrypting messages with one-time pads; and a page on processing analogue art. Their tooling, uxn, is a small virtual machine that has accumulated a real following as a demonstration of how much is possible on trivially portable hardware.

The through-line of their practice is not minimalism for its own sake, it is comprehensibility. One commenter quotes a log entry where Bellum unpicked an old hat to work out how to make a new one and compared clothing to open-source software. That is the actual thesis: not that off-grid life is better, but that you should be able to understand, repair and recreate the things you depend on, and that this is a discipline rather than an aesthetic. The cryptography pages follow from the same instinct — direct action organising happens on platforms that spy on participants, so teach the friends how to encrypt instead of complaining about the platform.

The best objection comes first in the thread and is worth taking seriously: living off-grid is fundamentally solipsist, and abandoning society rather than trying to live well inside it is a dead end. The reply is implicit in everything they publish — they are not off-grid for themselves, they are off-grid and then document the entire method for free, including the parts that fail. That is not withdrawal; it is a public experiment with the results posted. Whether it scales is a separate question, and the honest answer is that it is not meant to.

15-year search for a band that charted once and vanished — 338 points

338 points · 123 comments · shahidhussain.com · HN discussion

In 2011 the author bought a CD at a gig with two tracks on it, “Electric” and “Searchlights”, by a band called Salvage. The recordings were obviously professionally produced and the band was otherwise untraceable. He kept looking, on and off, for fifteen years.

He found them. The write-up is an interview with the members and their producer, and the payoff is better than a reunion story. The tracks were recorded in Cardiff at Manic Street Preachers’ FASTER studios on the vintage Trident desk from Rockfield, using the Manics’ own instruments and amps, produced by someone who turned up to their rehearsal room first to hear how they played live. Every label passed. One senior A&R explained why, and it is the most quotable line of the day: the song was “something a band would drop on a 3rd album, not the 1st. It’s too mature.” The band changed style, got UK industry attention under a different name, played Glastonbury, still never got a deal, and wound up in late 2008. The members now write for TV and film as Glacier and Summit.

The reason the story lands is that it documents a failure mode nobody writes down: not a band that was bad, and not a band that was unlucky, but a band whose best material was rejected for being too good too early. The thread is full of the same genre — the Reply All episode about the missing hit, a Björk track in Icelandic, an obscure cow-punk single from the eighties — and the recurring technical cause is streaming. When your music library was local files you formed spatial and contextual memories of it; when it is an infinite rented catalogue with no track names, the thing you are trying to remember has no handle to grab.

334 points · 204 comments · arxiv.org · HN discussion

Bastounis, Circelli and Hansen attack the verification story from the one direction nobody had checked. Autoformalisation — an AI translating a natural-language proof into Lean so a kernel can check it — is being treated as the trust mechanism for AI mathematics, including OpenAI’s announced blow-up result for Navier–Stokes. Their claim is that the formalised Lean proof does not correspond to the natural-language proof: the translation silently altered the argument.

Read the disclaimer before the headline, because the paper is careful and the coverage has not been. The authors explicitly do not claim OpenAI’s natural-language proof is wrong. They claim the Lean artifact does not represent it — and that a machine-checked certificate therefore certifies something other than what the release page pointed at. The thread’s strongest defence of OpenAI is that natural language is imprecise enough that multiple Lean translations are all legitimate, and that if the top-level theorem statement matches the Clay Institute formulation, the intermediate mismatch is of no consequence. That is a real argument, and it depends entirely on an assumption the commenter concedes is non-trivial: that the formalised statement is the one mathematicians actually meant.

Which is the generalisable finding. A proof checker verifies that a derivation follows from axioms. It does not verify that the axioms encode your intent, and it says nothing about whether the prose you shipped describes the derivation that was checked. Lean closes the gap between a statement and its consequences. It does not close the gap between a statement and what you wanted to say — and that second gap is now the load-bearing part of every “the proof is machine-checked” claim, including the ones in this roundup’s other maths entries.

Anti-patterns in software blogging — 316 points

316 points · 146 comments · refactoringenglish.com · HN discussion

A catalogue of the ways technical posts fail: the meandering intro (by far the most common), the assumption that the reader knows everything you know except the one thing you are writing about, overreliance on links, the “sequel injection” where a post becomes an excuse to announce the next post, excessive formality, and mundane rendering failures — mobile overflow, unreadable fonts, broken HTML.

The prescription is uncontroversial and correct: state the point first, give the reader a familiar anchor before the unfamiliar material, and stop building suspense in technical writing. The best comment sharpens it into a rule no style guide contains — education is not storytelling, and withholding the conclusion to preserve a reveal actively harms comprehension. A good presentation tells you at the start what it is going to say, then says it, then says it again. The second-best comment is a quote worth stealing: the sole purpose of the first sentence is to get you to read the second.

The thread also contains the complaint that everyone is thinking. A commenter estimates that the majority of software blogs encountered now read as LLM-written and are “usually terrible,” and asks whether anyone can just tell the models about these anti-patterns. They can. It does not help much, because the failure mode of generated prose is not ignorance of the rules — it is the absence of the thing the rules were formalising, which is having something specific you wanted to say. The list of anti-patterns is a description of what writing looks like when nobody had a point.

I gave Opus 5.5 one prompt and six hours to visualize Invisible Cities — 305 points

305 points · 158 comments · quesma.com · HN discussion

One prompt, six hours of agent time, and a generated interactive visualisation of all 55 cities in Calvino’s Invisible Cities, with each city rendered as a scene rather than a diagram. The author is an experienced hand at this — he has done data visualisation and explorable explanations work through Opus 4.8, Fable 5.1 and GPT-6 Astra, and describes 5.5 as the point where AI-assisted design stopped requiring him to strip out slop and fix overlaps by hand.

The thread’s reaction is far more interesting than the artifact, and it splits along a line that has nothing to do with quality. The people who have read the book mostly did not want the thing to exist: part of the book’s function is that you conjure the cities yourself, and a rendered version becomes the memory you keep. One commenter clicked a city at random, found the opening phrase was about rejoicing at each different bridge, and counted five bridges total, three of which did not span anything and instead ran along the current. That is the failure mode of visualisation-as-interpretation in a single example.

The second reaction is harder to argue with. “I understand why that’s impressive. But I don’t feel much.” It reads like a well-typographed deck somebody made ten minutes before the meeting. And that is the actual shift this launch demonstrates: the ceiling of one-shot generative work is now high enough that the output is competent and the interesting question has moved from can a model do this to why would anyone want the version a model produced. The answer might be that most people do — most slideshows, most dashboards, most landing pages are labour, not craft, and should be generated. The mistake is assuming the same holds for the things people make for love.

Docker Agent — 292 points

292 points · 134 comments · github.com/docker · HN discussion

Docker’s entry is a docker CLI plugin that runs agents defined in declarative YAML, with a tool ecosystem and multi-agent orchestration, written in Go. It ships pre-installed in Docker Desktop 4.63+, is available via Homebrew, and the pitch is “no code required” to build and run agents that collaborate.

The interesting content is the reaction rather than the product, because it is a clean read on where the space is. The top-voted criticism is that “no code required” stopped being a selling point the moment everyone had an agent that can emit mostly-correct code on demand — the value of a no-code layer was always the escape from writing boilerplate, and the boilerplate is now free. Two commenters observe that agent harnesses have become what JavaScript frameworks were in the 2010s: every vendor has one, they are all converging on the same shape, and the differences are mostly naming. A third lists five competing projects (kagent, Kubernetes’ agent-sandbox, LangChain’s deepagents sandboxes, and others) and asks outright what each is for, which nobody answers.

There is one substantive technical complaint and it is the right one. Someone went looking for security documentation and had to dig for it, then notes in the same breath that Docker’s terms now permit collecting agent inputs and outputs. For a category whose entire pitch is letting a model execute tools on your behalf, the configuration that matters is the sandbox boundary and the data retention clause, and neither is on the front page. The recurring theme in the thread — orchestration is not the hard problem, agent drift over long horizons is — is the more useful framing, and it is the thing none of these harnesses currently solves.

How did Rosalind Franklin miss the helix in her iconic DNA image? She didn’t — 287 points

287 points · 97 comments · science.org · HN discussion

A science historian and an X-ray crystallographer — Alistair Sponsel of the Science History Institute and Brian Sutton of King’s College London — re-examined Franklin’s notes on an earlier diffraction image, photograph 49, made days before the famous photograph 51 from the same DNA sample. Their argument is that the earlier image had already shown her the helix, that this is precisely why she made a second, more refined version, and that the epiphany Watson later claimed for himself was not his to claim.

The pushback in the thread is more careful than the article, and mostly correct. Helical DNA was already the leading candidate by 1951 on the strength of cruder diffraction images that had existed for a decade, so “Franklin knew it was helical” is a much smaller claim than the framing implies. The thread decomposes the discovery into three separate questions — that DNA is helical, that it is a double helix, and that the two strands carry complementary bases — and assigns them to different people, with Franklin and Wilkins settled on the first and Watson and Crick on the third. The contested ground is only the middle one. The published paper’s own abstract concedes it refutes a specific claim about Franklin’s failure, not that she discovered the structure. Several commenters also flag the headline for bait, which is fair, and one notes that HN’s original submission title asserted something the article never says.

The most useful point is the least dramatic: the abstract and the headline are answering different questions, and only the abstract is defensible. The reason this matters beyond one historical dispute is that the thread’s instinct — read the paper, not the framing — is exactly the instinct the mathematics entries above are asking for.

Whistle: Speech to Text in 16.9 MB — 273 points

273 points · 70 comments · cactuscompute.com · HN discussion

A speech recognition model shipped as one 16.9 MB file, running on CPU with no dependencies, loading into the same C++ engine as the company’s Needle model so a single binary can go from audio clip to tool call. It covers English, German, French, Spanish, Italian, Dutch and Polish, handles up to 30 seconds per utterance, reaches first token in 11 milliseconds, and processes audio entirely on-device.

The architecture disclosure is unusually complete — sample rate, mel bins, window and hop, the convolutional stem’s channel counts and downsampling factors are all in the post, which is more than most model announcements provide and makes the size claim checkable rather than asserted. If you are embedding recognition in a wearable or a microcontroller, 16.9 MB with no Python and no network is a genuinely different product from a 1 GB Whisper checkpoint, and the seven-language coverage matches the plausible European market for exactly that hardware.

The thread tests the claim and finds the honest limits. There is no streaming output in the demo — you record, then wait — which one commenter correctly calls essential for live transcription and which the release does not advertise. A commenter running it over a TV episode hit a repetition failure where it emitted “Thank you” for sixty seconds of unrelated dialogue, which is the standard degeneration mode for small decoders and worth knowing before shipping it. And the sharpest comment is the framing one: the hard problem in speech recognition was never binary size, it is understanding an 84-year-old Croatian stroke survivor with a drooping mouth who is trying to dictate an autobiography. That user does not need a smaller model. They need a much better one, and 16.9 MB is not going to get there.

‘Jonathan’ is the oldest land animal on Earth — 263 points

263 points · 152 comments · 404media.co · HN discussion

Jonathan is an Aldabra giant tortoise on St Helena, roughly 194 years old, hatched before the American Civil War and photographed in the 1880s. Researchers have now sequenced his genome for the first time, and report unique variants in most aging pathways — DNA repair, insulin regulation, mitochondrial function — plus unusually low-entropy DNA methylation.

The reporting is careful and the underlying result is a single genome, which is the thing to keep in view. Every longevity genome study returns variants in the same pathway list, because those are the pathways that govern ageing in every animal examined; finding them in a 194-year-old tortoise is confirmation that the machinery is conserved, not an explanation of why this particular tortoise is old. The genuinely interesting comparison is in the thread rather than the paper: no land animal has cleared 200 years, while several marine species run to five or six hundred. If extreme longevity were a solved molecular problem, the ocean would not be where the record holders live — which points at metabolic rate, temperature and predation pressure as much as at gene variants.

The headline leap to human application (“so we can live longer too”) is the press framing and not the study. The best joke in the thread is that min-maxing longevity is a build trade-off against speed and charisma, which is both funny and roughly how comparative biology actually works.

OpenAI annualised revenues $20B less than previously signalled — 250 points

250 points · 147 comments · cnbc.com · HN discussion

OpenAI told investors it hit roughly $50 billion in annualised revenue at the end of September. The figure widely reported late last month was $68 billion. Nvidia, Oracle and CoreWeave sold off on the difference. A person familiar with the matter told CNBC the higher number included gross revenue from OpenAI’s partners, and was presented that way to make comparison with Anthropic more direct.

That explanation is the story, because it is an admission that a number circulating as OpenAI’s revenue for weeks was never OpenAI’s revenue. Partner gross revenue is a different quantity with a different meaning: it counts money that flows through a partner’s books, which is why it is a flattering number for a company being compared against a rival with different partnership structures. The difference is not an accounting error, it is a definition substitution, and it was performed in the direction that made the number look bigger. Per the investor presentation, July alone was close to $30 billion annualised, which puts the trajectory and the revision in the same sentence for anyone who wants to do the arithmetic.

The thread is mostly doing that arithmetic and drawing the conclusion the market did. The most quotable comment is the simplest: does any other company get to misreport $20 billion in revenue and still be taken seriously. The runner-up is a fair hit on the media rather than the company — the same outlets printed the high number to make one story and the low number to make another, and the second story’s news value is largely that they printed the first one. The structural point underneath, from the top comment, is that a company whose valuation supports a meaningful fraction of the entire index is disclosing financials on a voluntary and negotiated basis, which is what happens when something systemically important stays private.

Cleo (Mathematician) — 248 points

248 points · 56 comments · wikipedia.org · HN discussion

Vladimir Reshetnikov posted 39 answers to hard integration problems on Mathematics Stack Exchange between roughly 2013 and 2015 under the name Cleo, each giving the exact closed form with no intermediate steps. The accuracy and the refusal to show work convinced a lot of people that Cleo was a collective or a system rather than a person. Reshetnikov revealed himself; a Reddit user’s forensic analysis of posting patterns in late 2023 is credited in the Wikipedia article as the confirmation.

The best technical observation in the thread is from a non-mathematician and it is correct: arbitrary-looking hard integrals are trivially fabricable by starting from a function with a known closed form and differentiating it, which means a good portion of the display was a magic trick, and the Wikipedia summary does not say so. That does not make the answers unimpressive — the point of a magic trick is that the audience cannot do it — but it does mean the answers had less content than they appeared to. It also raises an unaddressed question a commenter asks directly: whether any of the 39 came from a question that was not plausibly from an accomplice.

The thread’s real subject is the unmasking, and it splits honestly. Contributing something good while maintaining anonymity does not entitle anyone to be left alone, but outing someone who was doing no harm is a rude thing to do and the sleuths should own that. The most interesting comment connects Cleo to today’s AI proofs: the answers were correct and useful, and the value of the work — showing how to get from the problem to the answer — was exactly what was withheld. Read that alongside Tao’s thread and the OpenAI retractions and it stops being an internet mystery story.

House with 15m underground tunnels for sale for 300k — 229 points

229 points · 213 comments · readingchronicle.co.uk · HN discussion

Tony Deane, a civil engineer, spent about twenty years hand-digging a network of rooms and corridors up to 15 metres below his house. The property is going to auction with a guide price around £300,000. One room has an entrance concealed inside a rocket-shaped structure by the pool.

This is the item on the page that needs no analysis and benefits from a single observation: the commenters immediately found the auction listing with floor plans and more photographs, confirmed the depth figure is metres and not millions of tunnels, and located two comparable cases — William Lyttle, the “Mole Man of Hackney,” and a TikTok creator documenting her own excavation. The genre is real, it is mostly Britain, and it is populated by people with the relevant professional qualification. One commenter calls it a Bond villain’s starter home. Another correctly points out that a 23-foot rocket is a terrible way to hide a secret entrance. Both are right.

OpenAI withdraws three mathematical results — 224 points

224 points · 518 comments · github.com/openai · HN discussion

The repository’s own changelog, dated October 7, is the primary source and it is worth reading in full because it is written like a software release. A sign error in “Algebraicity of Weil classes on split abelian eightfolds” invalidates a stabilisation-trace cancellation argument used by two dependent papers, so three manuscripts were withdrawn — the Weil classes paper, the Kuga–Satake correspondences result for K3 surfaces, and the rational Hodge conjecture for products of K3 surfaces. Fourteen other manuscripts were revised for proof repairs, corrected statements, clearer hypotheses and dependencies, and one obsolete citation. Withdrawn papers carry notices explaining the gap and linking the archived version.

Three errors in a first pass over roughly 400 results in twenty-four hours is a good error rate and a bad signal at the same time, and both halves are worth saying. Good, because a sign error caught by the community within a day of publication is precisely the verification process working, and because the repository documented the retraction rather than quietly replacing the file. Bad, because the three papers were not independent — one invalid argument took out two results that depended on it, which means the failure mode is correlated and the true error rate is not 3/400. And bad for a third reason that the thread names immediately: the release was accompanied by an explicit request that the community carry the verification work, and the first thing the community’s verification found was a sign error in a dependency chain. The same commenter notes that six new papers were added in the same window, which the headline does not mention.

The framing the thread settles on is the sharpest thing anyone has written about machine mathematics this week: this is what it looks like when software engineering practice meets mathematics. Version-numbered releases, papers withdrawn as a batch, a changelog tracking repairs, and a public dependency graph by implication. One commenter writes the parody changelog — “release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47” — and the joke lands because that is genuinely the workflow. The open question in the thread has no answer and is the right one: whether the withdrawal came from a human mathematician reading the paper or from an automated consistency check, because those are two very different levels of assurance and only one of them scales to a corpus this size.

The Slow Formation of Durable Software — 222 points

222 points · 98 comments · newsletter.dancohen.org · HN discussion

Dan Cohen tells the origin story of Zotero, the citation manager used by more than twenty million people, from the first alpha in June 2006 — back when it was called SmartFox, renamed after consulting both a lawyer and an Albanian dictionary — to the decades of funded work behind it. Five years from conception to a rough prototype, and years more after that. The argument he draws is that this could not have been accelerated: “we did not know exactly what we wanted, and so could not have written coherent prompts for an LLM.”

The comment thread splits exactly where the argument is strongest and weakest, and it is a clean split. Nobody really disputes the conception claim — a growth and acquisition person makes it better than Cohen does, pointing out that the product’s value was the slow discovery of what it should do, and that this discovery is the deliverable rather than a delay in producing it. Where Cohen gets pushback is execution: one commenter who works on features at thousand-engineer companies says “software can be created more or less instantly with AI” is simply false at scale, where a feature involves five teams and a year of coordination; another says LLMs are in fact excellent precisely when you do not know what you want, because you can iterate against a running prototype instead of against a prompt. The romanticising charge — that LLMs would have turned five years into five months — is asserted more confidently than it can be supported, but it is not absurd.

The most useful framing in the thread is the three-motion loop: deciding what it should do, making it do that, and checking that it actually does. AI collapses the middle motion and leaves the first and third, which is why the experience of building with it oscillates between startling speed and complete stasis. Cohen’s post is a good argument about motion one. It reads as a claim about all three, and that is where the disagreement comes from.

God of War on PSP, recompiled to WebAssembly and running in the browser — 221 points

221 points · 110 comments · github.com · HN discussion

This is not an emulator. The PSP’s MIPS machine code is statically translated ahead of time into C++, compiled to WebAssembly, and linked against a small reimplementation of the console’s operating system and graphics chip that draws through WebGL2. God of War: Chains of Olympus boots through menus, cutscenes and combat at a measured 60 frames per second in Chrome and Firefox at up to four times the PSP’s native resolution, with music, speech and sound effects intact; video playback is skipped. Ghost of Sparta followed through the same pipeline, and there are on-screen touch controls for phones.

The significance is in the technique rather than the game. Emulation has to reproduce a machine’s behaviour at runtime, which costs an interpreter layer or a JIT and a great deal of per-title work; static recompilation instead converts the original binary once and ships native code, which is why the frame rates here are not achieved by a general-purpose emulator. The comparison in the thread makes the point: someone ran this game on an i7 laptop a decade ago at 15 fps. The generalisable consequence is that a console with a well-understood instruction set and a small OS surface is now a tractable target for a single determined person plus tooling, which is a change in who gets to do preservation at all.

The obvious caveat is the legal one, and the thread raises it in the first few comments: the games are still under copyright, the recompilation ships without them, and nobody believes Sony will leave this alone. The subtler point made in the thread is that this is why preservation needs a legal theory — the publishers are not making money from these titles, they are not re-releasing them, and they will still send the letter.

How machines learned precision — 218 points

218 points · 68 comments · glinscott.github.io · HN discussion

An illustrated history of the bootstrap problem in machine tools, built around Henry Maudslay’s screw-cutting lathe in London around 1800. The problem is circular and it is stated cleanly: a lathe cannot cut an accurate screw unless its own rails are straight and its own leadscrew is even, and in the 1770s workshops could not reliably make either. Errors compound rather than cancel — a bent rail makes the tool follow the bend, an uneven screw cuts an uneven thread, and the result is leaking pistons, wobbling shafts and jammed nuts.

The techniques that broke the circle are the payoff. Maudslay’s trick for making a screw without a reference screw: use two imperfect screws to guide the cutting of a third, geared to turn at the same speed, with a bar joining their two nuts and the tool fixed at the midpoint so its position is the average of the two — which cancels the error common to both. After repeated copying and correction he produced a master screw five feet long, two inches across, fifty threads to the inch, with a foot-long nut engaging six hundred threads at once. That is one of the great bootstrapping results in engineering history, and the article’s animations make the geometry legible in a way prose does not.

The thread is unusually good on the human side. A commenter whose father ran a machine tool shop in northern India before CNC wiped out demand for manual lathes describes what happened to a whole class of skilled work; the author shows up to say it started as research into Watt’s cylinder-bore problem for an earlier piece on beam engines and spiralled out. The best structural observation is that this is the same problem AI is now in with data: you need a source of quality to produce quality, and there is no external reference to check against until you have built one. The Maudslay answer — average two independent flawed sources and iterate — is not the worst analogy for how synthetic data generation works either.

Why were Victorian elites so effective? — 213 points

213 points · 368 comments · worksinprogress.co · HN discussion

The article runs an experiment in counter-programming. Modern success writing prescribes austere routines — 5am starts, cold baths, no alcohol, obsessive hours, STEM-heavy schooling. The Victorian elite who ran the largest empire in history did the opposite: late risings, litres of port, schools and universities that were barely academic and mostly devoid of science and engineering, and an enormous amount of unstructured social time. The argument is that these were not vices the elite overcame but mechanisms that produced the outcomes the modern template optimises against — networks, judgement, long-horizon relationships, and the bandwidth to absorb a great deal of information without needing to act on it immediately.

The thread’s best objection is quantitative and it lands: the governing apparatus was vanishingly small. The colonial office in London had fourteen staff in 1816, growing to a few dozen by mid-century, which means the empire was run by an elite small enough that coordination happened through personal acquaintance. When your decision-making body is fourteen people, unstructured social time is the coordination mechanism, and the modern template optimises for a world where organisations are thousands of people and coordination needs process instead. The vices were not virtues, they were the shape of a network.

The second objection is the one that should stop the article cold, and it is stated bluntly at the top of the thread: the Victorian elite were optimised for rent-seeking on an economy and an empire that other people built and worked. Their effectiveness, judged on their own terms, consisted of extracting surplus from a system they did not create. That is a genuine reframing rather than a nitpick, and the article never engages with it — which is the problem with reading history as a management manual. The interesting question is not what routine produces effective elites, it is which of their advantages were merit and which were monopoly.

Beauty in DVD Menus — 206 points

206 points · 130 comments · vale.rocks · HN discussion

A tour of DVD menus as a design medium, from the format’s 1996 launch through its 2000s peak. The article is precise about the constraints that shaped the aesthetic: NTSC video at 720×480 (PAL at 720×576), a limited number of subtitle and audio tracks doubling as button highlight layers, and enough storage headroom that the designers of a big release could afford full motion video behind the navigation. The result was a period of genuinely creative interface work — animated menus, hidden features, navigation games — that the Blu-ray era mostly flattened into a video clip with an overlay.

The thread does the two things you want. First, it supplies the primary sources: a commenter has archived roughly 250 DVD menus on disk, noting that some run over a gigabyte because the menus are themselves video that must play to function, and another names Memento’s hidden button combination that played the film in reverse, scene by scene. Second, it supplies the counterargument, which is stronger than the nostalgia. One commenter describes the DVD menu as reverse-engineered anti-feature — by the time you have navigated the animated menus and been shown every memorable shot in the film, you no longer want to watch it. Another’s version is worse: the elaborate menus arrived on a physical medium whose remote controls were bad, whose players were slow, and whose every button press took two seconds to register.

Both positions survive, and the resolution is that the design deserved its audience. Enthusiast releases got elaborate menus and the enthusiast audience wanted them; the mass-market discs that ship today are plain because most people want the film and nothing else. That is not a decline in craft, it is a product decision made by people who correctly read their users — and it is the same tradeoff every interface that tries to be interesting eventually loses.

Still on the page

Fifteen of the stories that cleared 200 points today were covered in earlier roundups. Deltas are against the last roundup that tracked each story, which for all but one is yesterday’s.

Two of the day’s biggest movers were yesterday’s launch announcements, which is where the points go once a release leaves the front page. Claude Haiku 5.5 went 360 → 1,027, up 667, and the thread’s useful content arrived with the second wave: the pricing is stranger than the headline, at $0.10 per million input tokens under 100,000 tokens and $0.50 above it, so the model is cheapest exactly where it is least useful and the cutoff does not apply to Sonnet or Opus. GPT‑6 and Intelligent UI went 256 → 740, up 484, and the criticisms in its thread are consistent enough to be worth reading as a review: elaborate generated layouts where a plain answer would do, unverifiable claims about the model’s own video player, and a pointed observation that the new “intelligent UI” is also a newly announced advertising surface.

Sharing AI progress in mathematics held on at 1,185 → 1,316, up 131, second on the page and the source of the retraction story above. Visa, Mastercard and the banks went 354 → 586, up 232, which is a strange amount of attention for a class action until you read the thread and find people pricing merchant discount rates against gas-station loyalty apps. JPEG XL in Chrome went 405 → 559, up 154, and Tell HN: GitHub refuses to remove cracked copies went 491 → 558, up 67. Smaller movers: the Commodore 64 keycap font 348 → 398, up 50; the Nobel Prize in Chemistry 262 → 297, up 35; EmbeddingGemma 2 411 → 429, up 18; Anthropic’s diary disclosure 822 → 831, up 9; the Nobel Prize in Physics 568 → 576, up 8; Strands Decider 2B 274 → 281, up 7; What’s Earth’s dominant species by mass? 290 → 296, up 6; and the GitHub incident 222 → 226, up 4.

The one item that has finally stalled is the one that has been at the top of the page for four days. Mistral Large 4 went 1,980 → 2,025, up 45, and its thread has turned into an argument about whether near-parity with the frontier is a triumph or a failure for a model trained on 3,800 GPUs in the company’s own European datacentres. Both readings are represented and neither is stupid, which is roughly the correct level of enthusiasm for a very good model that is not the best one. Worth recording alongside it: there is a second submission of the same launch, “Mistral Large 4: Le Chonk,” which has now accumulated points without anyone covering it, and a duplicate submission of the OpenAI withdrawal story hosting 337 points of its own, 113 more than the version linked above.

Throughline

First: the verification layer everyone has been pointing at is thinner than advertised, and today three separate stories established it in three separate ways. The Navier–Stokes paper shows that a Lean certificate can be a faithful derivation of a statement that is not the one the prose described — autoformalisation moves the trust problem rather than solving it. The OpenAI changelog shows what happens when you run a research group like a software release: a single sign error propagates into two dependent papers and the retraction notice arrives twenty-four hours later, which is fast and also tells you the failure was correlated rather than random. And Aaronson’s own thread supplies the denominator nobody quotes — the model was tried on about 8,000 problems and solved roughly 5% in a single three-hour attempt each, which is a staggering capability and a long way from the impression that mathematics has been wrapped up. Tao’s contribution is the one that will still matter in a year, because he is not arguing about capability at all. He is pointing out that a discipline whose entire prestige economy was priced against being first to solve has just had its core activity commoditised, and that no mechanism exists for the follow-up work — the exposition, the community, the new directions — that made solutions useful in the first place.

Second: the AI bill landed, and it arrived in the same week as the revenue revision. Microsoft cutting per-employee AI budgets from $100,000 a month to $10,000 and OpenAI’s annualised revenue turning out to be $50 billion rather than the $68 billion that circulated for weeks are the same story told from opposite ends. The $100,000 monthly allowance was a policy decision made when the models were a strategic hedge and nobody was counting; $10,000 is what the number looks like when someone finally is. The revision to the revenue figure is the part that deserves more attention than it got: the higher number counted gross revenue from partners, and the explanation offered was that this made comparison with a competitor more direct, which is an admission that a number presented as OpenAI’s revenue never was. Meanwhile the Victorian elite thread supplies the day’s cheapest lesson in scepticism about a number that looks impressive on its own: an empire covering a quarter of the world turns out to have been administered centrally by fourteen people, which is not a claim about how much work sixteen people can do. It is a claim about how much of what gets counted is really being done by the part of the system nobody is counting. The same question applies to a partner-gross revenue figure and to a corpus of 400 machine-generated proofs.

Third: the day’s non-AI half is about the same thing as its AI half, which it usually is when you look. The recompilation of God of War is a static-analysis win over a decade-old binary; Maudslay’s lathe is the bootstrap problem, where you have to derive accurate reference from inaccurate copies by averaging and iterating; DVD menus are a design medium killed by the correct observation that most users want the content and not the interface; Hundred Rabbits is the discipline of being able to repair the things you depend on; and the fifteen-year search for Salvage is a man who kept a physical artifact and used it as a handle to recover everything around it. Cohen’s Zotero post makes the link explicit — you cannot write a coherent prompt for a product whose requirements you have not yet discovered, and the slow conversation in which the requirements emerge is the work, not the overhead. Each of these is an argument that the artifact matters and the process that produced it is load-bearing. That is not a rejection of generated work; it is a specification for where generated work does and does not fit.