Forty stories cleared 200 points today, and twenty-two of them are new — the biggest new crop in a week, because yesterday’s page was mostly recycled material aging upward. The top two are still the e-ink bird frame at 2,385 and the AI-poster essay at 1,849, both submitted days ago. The Genuinely New split roughly into three piles: archives that nobody wants to be responsible for, models whose published numbers no longer describe what you get when you call them, and people doing skilled work that the economics has stopped paying for.
What happened to the Snowden archive
653 points · libroot.org
The last document published from the Snowden archive was released on 29 May 2019 by The Intercept. The Guardian stopped in February 2014, Der Spiegel in January 2015, the New York Times and ProPublica in August 2015, and The Intercept closed its archive in March 2019 — the internal email from Michael Bloom framing it as resources and “editorial priorities” while noting the archive’s remaining value and hoping Greenwald and Poitras would find an academic partner. They didn’t. No outlet, journalist, or institution has published a document from the archive since.
The scale numbers are the part that should sit badly. The Guardian held roughly 58,000 documents and published about 30; the partner group’s combined output is what Rusbridger was describing when he cited 1% (in the same parliamentary session he also gave 26 of 58,000+ for his own paper’s output, which is 0.045%). The Intercept is believed to have destroyed its copy — that claim rests on an anonymous source quoted in Appelbaum’s 2022 doctoral thesis, and the Intercept has declined to answer the question twice, then declined to answer this article’s questions along with First Look Institute and eight former staffers. The article’s authors spent August and September 2026 contacting more than twenty people and organisations. Two replied, neither with anything of substance.
Three people still hold complete copies: Glenn Greenwald, Laura Poitras, and Barton Gellman, who has the 50,000+ files Snowden sent him directly. None of them answers to a budget, a board, a funder, or an editor — the constraints the article can verify are gone. Poitras said in 2022 that the archive still exists and there is “a vast amount of information that hasn’t been reported of enormous contemporary and historical significance.” Gellman agreed it would be valuable for research, then said the opsec of sharing it is too hard: “I just put the whole thing in cold storage. I feel bad about that.” Greenwald has described it as hundreds of thousands of documents. None of the three has published a document from it in seven years.
The explanation that gets offered depends on who is being asked, and the accounts don’t reconcile. Greenwald in 2019 said the closure partly reflected his own intent to seek other partners; in 2021 he told New York Magazine that nobody at The Intercept wanted it closed and it happened because of a disagreement between Betsy Reed and Poitras, two years after he’d released documents implicitly saying he and Reed decided it. The strongest structural point in the piece is the one about why nobody will hand it over: publishing requires cold storage, air-gapped machines, and close reading, and the political economy of research institutions punishes taking custody — universities fear grant exposure and foreign-collaboration optics, which is exactly the reason Snowden gave for why the intended handoff to academia never happened. What’s left is a distributed archive whose every remaining holder has both the technical ability and the motive to hold it, and an Overton window that has already absorbed most of what it contains.
AX — Google’s Open Agentic Orchestrator
625 points · agentexecutor.io
A declarative orchestrator for agent workloads that runs on a separate runtime called Agent Substrate. Four primitives, expressed as ax.io/v1alpha1 manifests: Task for an isolated sandbox with CPU and memory limits and suspend/resume; Workspace to pre-wire Git repos, MCP servers and skill packages so an agent starts warm — or just describe the environment in English and let the platform build it; Gateway for an outbound allowlist that injects credentials rather than handing them to the agent; Model for one place to configure models, parameters and secrets out of a Kubernetes secret. The commands are the Kubernetes verbs: ax apply, ax watch, ax ssh. Claims: billions of tasks per cluster via lightweight actors, sub-second resumption of suspended agents, and dense multiplexing so idle waiting time becomes spare capacity.
The README opens with a warning that core concepts and protocols are still being refined and major breaking changes are expected before a stable release. The quickstart then says you need a Kubernetes cluster, ko, a container registry your cluster can pull from, and a reachable Agent Substrate control API. That gap between “uncompromising focus on ergonomics” and “install a Kubernetes operator” is what the thread noticed first, and it’s a real one — the density and resumption claims only pay off at a scale where you already have someone who operates clusters.
The more interesting objection is structural rather than ergonomic: an agent is a call to an external service that happens to be slow, stateful, and expensive, and the argument for flipping the system so the agent is the process is that the existing paradigms don’t have a good story for long-lived state plus untrusted code plus per-task credentialing. That’s a genuine gap. But it’s also how Kubernetes itself got sold, and the “billions of tasks” number is doing marketing work — nobody has billions of agent sessions, and the design choices that matter (how a Gateway allowlist fails closed, what a suspended actor’s checkpoint actually contains, who audits the credential injection) aren’t in the pitch.
Samsung is expected to more than double HBM4 and HBM4E output next year
544 points · Seoul Economic Daily
The lead indicator here is glass. Glass carriers are temporary supports bonded to the underside of an HBM wafer so it doesn’t bend or crack while being ground thin and drilled, and they’re used repeatedly after cleaning. Samsung’s outsourced cleaning volume for them goes from 10,000 sheets a month last year to 20,000 this year to 50,000 next year — two and a half times current volume. Since carriers are reused and consumption varies with process loading and yield, the analytical move is an inference from a consumable rather than a disclosure from the company, which is why the article hedges properly: analysts conclude HBM4-family output will grow at least twofold. The rest of the numbers point the same way. Total HBM wafer input goes from roughly 180,000 to 250,000 wafers a month, about 40% growth, while HBM4-family shipments move from roughly 40% of the mix to about 80% as HBM4E ramps. Samsung shipped HBM4 in mass production in February with 1c DRAM and a 4nm base die, and provided 12-layer HBM4E samples in May.
Read it as a supply-chain signal rather than an AI-demand signal. The thinning step that makes HBM possible — grind the wafer to nothing, drill through-silicon vias, stack, and control warpage the whole way — is the unglamorous manufacturing constraint that the coverage usually skips, and it’s also where the sanctions bite. A commenter’s point is the one worth keeping: the credible bottleneck on Chinese accelerator output is HBM capacity at CXMT, not processor dies and not ASML. You can compensate for bad DUV yields with more wafers and smaller dies. You cannot compensate for not being able to stack the memory.
Anyone who buys DRAM for anything other than HBM should read the same numbers as bad news. Every wafer that goes to HBM4 is a wafer not making consumer memory, at a moment when the memory market has already been repriced.
Spain orders blocks on Archive.today and its mirrors
522 points · Reclaim The Net
Spain’s Second Section of the Intellectual Property Commission — part of the Ministry of Culture, not a court — ordered the blocking of several Archive.today domains and mirrors. No ruling, no hearing described in the coverage. Internet users in Spain who try to reach the site get redirected to a government page headed “ESTÁ USTED INTENTANDO ACCEDER A UN SITIO WEB ILEGAL” that accuses the visitor of “illegally facilitating access to content protected by intellectual property rights” and warns that they are “contributing to an illegal and criminal activity” and endangering their own security. The accusation is aimed at the person typing a URL.
Spain having an administrative body with fast-track domain blocking is not new, and neither is the region using it primarily for football. Commenters from southern Europe describe the same pattern in Italy, France, Portugal and, partially, the UK; one person in Spain notes that non-technical friends already buy grey-market “devices” from local installers for a few euros a month, which tells you the enforcement is a tax on people who don’t know a VPN is the answer. Another reports that Cloudflare edge IPs get blocked during matches, which is how a rights enforcement measure turns into intermittent outages for unrelated sites, and a third reports that Starlink hasn’t joined the blocking yet. And then the commentary collapses into the obvious double standard: an archive service that preserves pages is a criminal facilitator, while a model that ingests the whole web for training is an industry. Neither position is coherent on its own terms, which is roughly the state of copyright enforcement in 2026.
Disney+: new user agreement allows ads before movies in all subscriptions
474 points · Consumer Rights Wiki
Section 2(k) of the Disney+ Subscriber Agreement says tiers described as “no ads” or “ad-free” are “generally free of commercial interruptions, with certain exceptions that may change from time to time.” The exceptions: where streaming rights or other limitations require content to play with ads, where ads are served in live or linear content and special events including replays, and “limited promotional content” — brief clips about the bundles including upgrade messages, plus branded content, product integrations and sponsorship messaging. The change dates to January 2025 and binds existing subscribers unless they cancel; in September 2026 Disney sent German subscribers an email styled as a “clarification” that ads can be placed before and after content on Standard and Premium tiers.
The thread split cleanly, and the split is the interesting part. One camp read the actual clause and argued it says what it says: embedded ads in live sports you can’t strip out, plus bundle promos, so the outrage is manufactured. The other camp pointed out that the distinction between an ad and a promotion for another tier of the same service is a definitional choice Disney made unilaterally, in a contract you accept by not cancelling, and that the trend line is what matters — the ad-free tier exists as a price segment, and the segment is only worth maintaining until someone models the revenue from breaking it. That last argument is the strongest one and it doesn’t depend on Disney: if you pay specifically to avoid advertising, you have identified yourself as a customer with disposable income and low tolerance for interruption, which is precisely the audience the advertising market prices highest.
Two caveats, because the sourcing here is weak. The wiki page carries a live banner saying it’s marked incomplete and reads as original research needing references, and the “all subscriptions” framing in the title is broader than the clause it quotes. The direction of travel is documented; the magnitude of the current change is not.
What Sun got wrong
394 points · Bryan Cantrill
Cantrill’s distilled answer, fifteen years after his 2011 version: Sun had become bored with the mechanics of running a business. The illustration is a 2005 startup running its infrastructure on OpenSolaris that wanted to buy Sun hardware — the vindication of open-sourcing Solaris, the model working exactly as designed — and could not get Sun to pick up the phone. When Sun did pick up the phone, it pitched the wrong product. Meanwhile the team filled out Dell’s web form at midnight and got Steve, the local Dell account executive, on the phone the next morning. Within two weeks they had pitched for a pricing tier, had servers in the datacenter, and had them leased on the company’s financials with no personal guarantees. The account is in a 2006 blog post called “The Sun Doesn’t Shine on Me.” Cantrill read it while squatting in an abandoned Sun office as Fishworks was getting started, and describes his heart sinking at a story of strategic success and operational failure together.
That’s the whole post, and it’s useful precisely because it refuses the nostalgia frame. The technical achievements that people remember Sun for were real and they didn’t save the company, because the part of the business that turns a willing buyer into a signed order is a boring operational muscle, and a company that has stopped exercising it cannot be rescued by a better strategy or a better product. Sun lasted a few more years. Cantrill left, and later joined that startup — Joyent — which is the kind of ending a company earns when it treats its own best customers as an inconvenience.
Attention is all you have
386 points · alicegg.tech
The Tetris effect — play long enough and you start seeing tetrominoes in clouds and half-asleep — as the frame for a complaint about feeds. What you focus on shapes your thinking, and increasingly the choice of what to focus on has been delegated to systems whose objective is session length. YouTube knows you like cooking and art streams and still slips in stock-market-bubble and war coverage because that’s what keeps you watching. Spotify’s algorithm inserts cheap generated filler between real tracks to reduce what it pays artists. LinkedIn buries the career news you came for under strangers’ promotion of whatever Microsoft is invested in. And the recommendation surface people used to consult for product opinions is now a mix of language models and trolls. The prescription is bookmarks: a few dozen sites you chose for a specific reason, which is a real distinction — a bookmark is an intention, a feed is an absence of one.
Where the argument is soft is the causation. Every platform named here optimises engagement, but the specific mechanisms are asserted rather than measured, and “the algorithm made me doomscroll” is doing some work for the author that the evidence in the post doesn’t cover. Still, the tetris-effect parallel holds up better than most tech-essay metaphors: the phenomenon being described requires no platform intent at all, only sustained exposure. Which is the uncomfortable version of the point — you can delete the apps and still be shaped by what you spent a decade looking at.
Grok 4.7
363 points · xAI
A new, larger base model than 4.6, with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours, better self-verification, and native training on the Grok Bot harness. Same price as 4.6: $2 per million input tokens, $6 per million output. The vendor’s own table has Grok 4.7 at xhigh at 46.3% on CursorBench 4.0, against 40.4% for Grok 4.6 high, 41.7% for GPT-5.6 Sol Max, and 51.8% for Fable 5.1 Max; 71.0% on DeepSWE v1.1 at high effort; 38.0% on Terminal-Bench 4.0 against Fable 5.1’s 57.9%; 19.6% on the Harvey Legal Agent Benchmark, where it leads the table by a wide margin; and 56.7% on HealthBench Professional, where it trails Fable and GPT-5.6. The framing is price-performance on long-running coding tasks, and that’s the one axis where the numbers hold up.
Two things about the release are more informative than the chart. First, a larger model at the same price means margin compression, and the release slipped roughly two weeks past its original date and landed the day before an Opus 5.5 was rumoured — a cadence that reads as reacting rather than leading. Second, the benchmark table is internally inconsistent in a way that independent testing picked up: one commenter measured reasoning-token counts across effort levels and found low and medium using similar totals while xhigh used fewer than high, and a direct-API rerun didn’t reproduce the pattern either. When effort level stops predicting token spend, the table’s xhigh column stops meaning what it says. That’s a measurement complaint, not a capability complaint, and it applies to every vendor’s footnotes from here on.
The subjective consensus in the thread is worth recording because it cuts against the numbers: Grok has been “just behind” for several months, and the gap between just behind and pushing the frontier has turned out to be wider than people estimated a year ago. On side-by-side image-to-HTML work, Astra’s output came out better polished and better composed; xAI’s price advantage is real, and so is the sense that price is currently its main claim.
Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
343 points · GitHub
Small decision models — 0.8B, 4B and 9B — built on Qwen3.5 following the architecture described in “Jev’s Architecture Unmasked,” with training code and frozen eval suites published under Apache-2.0. The API matches TypeSafe’s System One endpoint, so their Python SDK can be pointed at a local server. One request carries multiple questions against the same input: yes/no (noul), multiple-choice (choice), and rating (score), each returning confidence and a full probability distribution. Questions share the input text and cannot read each other’s answers, which is the design decision that keeps the answers independent. The README’s example routes a support ticket to returns/shipping/billing, flags escalation, and scores frustration in one pass: 101 input tokens, 161 output tokens, under 500ms, with escalate: 0.93 and a frustration distribution across the three levels. CUDA, ROCm, and Apple Silicon; the 4B and 9B fit in bf16 on a 32GB Mac.
The scepticism this deserves is specific and comes from inside the thread. If the job is classification over a fixed set of labels and you can supply 50 to 100 labelled examples, an embeddings model plus logistic regression gets about 95% on email triage, trains in under five minutes on a CPU, weighs less than a megabyte, retrains on-device, and sends nothing anywhere. That baseline is not a strawman — it’s a real competitor to a 4B parameter model whose weights you still have to serve, and its privacy properties are strictly better. What the Kev-class models add is a single request that answers heterogeneous question types with calibrated uncertainty, which is a genuinely different offering from a classifier. Whether that’s worth a checkpoint is an empirical question, and the second objection is that the answers so far are unverified: the API-shaped compatibility with a proprietary system, a Qwen/RLHF base claimed to be “Jev-like” while Jev reportedly trains differently, and a broader distrust of TypeSafe’s data retention policy with no zero-retention option offered. There is a boomlet of these models and very little of it has been evaluated by anyone who doesn’t ship one.
Grim Fandango Puzzle Document (1996)
343 points · The Gameshelf (mirror)
A 72-page first-draft puzzle design document dated 30 April 1996, credited to Tim Schafer with Peter Tsacle, Eric Ingerson, Bret Mogilefsky and Peter Chan, written to pitch Grim Fandango to LucasArts. It contains the structure, early artwork, background and solutions for more than 80 puzzles, plus content cut for schedule: the Pizza Demon, the Giraffe Lady, Bernard, the Dillopede, and a five-puzzle action climax with Hector LeMans. Schafer surfaced it in 2008 for the game’s tenth anniversary with a note that puzzles “were nuts. Obscure. Mean, even,” then took his post and the document down without explanation. The Gameshelf rehosted it for historic preservation, which is the only reason the link still works.
The two things people kept quoting are the funniest available evidence about how internal documents changed. Schafer admits the final puzzle wasn’t designed when the document was due, so he wrote two nonsense paragraphs and overlapped them in the file to look like a print formatting error had obscured it. And the document has a small box near the end with the instruction to restrict your fallen tears of joy to it. The thread’s read is the right one: that kind of personality in a design document is exactly what a 2026 planning process would classify as inefficiency, and the artifact that survived thirty years is the one with the jokes in it.
ZuckOff is a free app that sees Meta glasses before they see you
340 points · WIRED Middle East
Paweł Szydłowski, a Polish developer, launched ZuckOff on iOS in early August and reached roughly 5,000 downloads in the first month, with about 1,000 on the Play Store. The mechanism is unglamorous: smart glasses advertise themselves over Bluetooth so their owner’s phone can find them, and those advertisement packets carry manufacturer IDs, service UUIDs and sometimes readable names. He bought Ray-Ban Meta, Oakley Meta and Snap Spectacles units, catalogued their signatures, and wrote matching logic — including 0x0D53 for Luxottica’s camera glasses, 0x058E for Meta wearables, 0x03C2 for Snap, and an Oculus service UUID. Matches come with confidence bands and the evidence shown, so you can disagree with a flag; proximity comes from signal strength with no direction, because direction would be invented.
The limits are the story and the developer states them plainly: presence is not recording, glasses advertise loudest while powering on or leaving the case and can go quiet once paired, detection depends on a manual catalogue that misses new hardware, and RF-dense rooms degrade it. Headsets are logged but not flagged on the reasoning that you can see a headset coming, and an alert is worth more when it’s reserved for the camera you can’t. Free tier is foreground scanning; background monitoring, alerts, extended history and CSV export are behind roughly $9.99 a year. The name is a jab at Zuckerberg and also the surname of a Boston University journalism professor who was not consulted. European regulators have reportedly asked questions.
The interesting part isn’t the app, it’s what it proves about the device category. Meta’s glasses were designed to connect to a phone automatically, and that convenience is implemented with an open broadcast any nearby device can read. Every detection system built on this will be one firmware change away from irrelevance, and the change that defeats it — stop advertising once paired — also degrades the pairing experience for the person wearing the glasses. That’s the actual lever, and nobody has pulled it yet.
I am often wrong
321 points · Boris Cherny
An internal note to his team, reposted. Six steps applied to essentially every problem: understand the information available, gather what’s missing, define the problem, define a clear and simple approach, define a goal, act with urgency. New information means going back and redefining three through five, repeatedly, and the churn is healthy if you know it’s the process rather than a failure. The named failure modes are the two most people skip: not defining the problem clearly, and not defining a simple enough approach. Both produce the same symptom — complicated plans with success criteria nobody can evaluate.
It’s short and it isn’t novel, and it holds up anyway because it’s written as a description of a working habit rather than a framework to adopt. The one line worth extracting is that the person making the plan is the worst-positioned person to notice the plan isn’t clear, which is why the feedback loop gets stated as an obligation rather than a virtue.
MCP was always a bad idea?
312 points · maharship.com
The argument: MCP shipped in November 2024 for models that couldn’t reliably call APIs, exploded into an ecosystem, and then models got good at exactly the work the protocol was compensating for. Context bloat from dozens of servers and their schemas forced harnesses to build search-and-execute shims (Composio, MintMCP, Pipedream) whose real job was credential centralisation. Now agents write scripts, compose services, and discover CLIs with --help, and most remote MCP servers are wrappers around HTTP APIs that already exist. The proposal is to standardise the thing underneath: agents sending Accept: text/markdown so servers return clean text instead of HTML or verbose JSON, and putting the preferred programming language in Accept-Language so documentation sites serve the right SDK examples — a suggestion from a Vercel engineer that Shopify has already implemented, which is the strongest evidence in the post that agent-oriented content negotiation is arriving without a new protocol.
The rebuttals in the thread are better than the post, and they name the thing the author skipped. Simon Willison’s list: control over exactly which external services an agent can reach, authentication that never hands the agent API keys, a UI for a human to connect and authorise services, and audit logging. None of those are tool-invocation problems, and a terminal agent with curl doesn’t solve any of them — it just moves the sandbox. A second rebuttal is that MCP is the same protocol for HTTP-based and stdio-local tools, which is exactly the case where talking about HTTP headers is meaningless. A third: the one-click plugin stores inside ChatGPT and Claude are MCP servers, and that’s what business users install.
Both halves are right about different layers. Direct API calls with content negotiation win for a coding agent in a terminal; MCP-shaped plumbing wins wherever a human has to authorise a service, an audit trail has to exist, or the tool never had an HTTP endpoint. The protocol probably folds into plain HTTP practice over time. The governance layer it was quietly providing won’t just disappear.
The Effect of CRTs on Pixel Art (2024)
301 points · datagubbe.se
A response to the viral argument that pixel artists used to design against CRT blur, so modern blocky pixel art is anachronistic nostalgia. The author agrees with the premise — CRT rendering did smooth low-resolution edges — and then points out the premise doesn’t explain much: the properties that mattered include signal quality, scanlines, the shadow mask, and the RGB cell structure, none of which a software filter reproduces convincingly. He uses real CRTs weekly and disables emulator CRT filters on principle. The post also carries an honest correction from 2024 conceding that the comparison screenshot he doubted was in fact a photo of a real tube.
The best counter-argument in the thread is that both sides are arguing about the wrong frame: pixel art is now its own aesthetic for high-DPI displays, not a failed imitation of a 1993 television. Read that way, the nostalgia isn’t misdirection — it’s a different medium borrowing a vocabulary. And the sharpest observation is economic: through the mid-1990s the best-paid artists in games worked in pixel art, so the constraint that produced the aesthetic was a budget, not a scanline. A commenter’s Atari 2600 example — tanks that look ridiculous at high resolution and plausible on a CRT — is the case for the other side, since the artists really did draw for the display they had.
Singapore’s National Library Board offers micropayments to build reading habits
279 points · Gadget Review
ReadSG launched on 6 September 2026 as a five-year national campaign. The mechanic: log at least 15 minutes of reading on GovTech’s CrowdTaskSG platform and earn 20 virtual coins, at a fixed 1,000 coins to S$1, with one qualifying session per day. That’s about two cents a day, and fifty consecutive days buys you one Singapore dollar. The headline writes itself and the thread immediately deflated it, noting that the actual programme is gamification with rewards bolted on — streaks, XP, leaderboards, prize draws, limited-edition items — where the money is a rounding error designed to be quotable.
So judge it as habit design, not incentive design, and it’s still worth watching. Two cents is too small to motivate anyone but large enough to appear in a headline and a press release, which is a second audience the programme is also serving. The interesting question is whether a government can buy a habit with streaks and a leaderboard, and the honest answer is that nobody knows at population scale, which is why the five-year window exists. The failure mode is the one every streak product has: people log fifteen minutes of a book they aren’t reading, and the metric tracks the logging rather than the reading.
Why do we need human mathematicians anymore?
275 points · Terence Tao’s blog, guest post by Po-Shen Loh
The context is real disruption: OpenAI announced a solution to the Millennium Prize variant of Navier–Stokes, and the response has been a wave of declarations — the Leiden Declaration past 4,000 signatories, Math and AI past 7,000, an open letter against the Caltech Mathathon past 2,000. Loh starts from the technologists’ objection that mathematicians should adapt and cede control, treats it as a proof to be checked, and finds a hole worth taking seriously. A useful detail given the subject: the prose is 100% his own, typed in vim with no generation, while the page design, layout and some headings were produced by Claude Code from his raw text file.
His chain: adopt the axiom that humans should help humanity flourish. Then note the observation that there are zero examples of a vastly more capable intelligent species surrendering decision control to a less capable one — Hinton’s baby-controlling-a-mother is the only partial exception anyone has produced — which means leadership has to stay human regardless of capability, because nobody has a proof that alignment will hold. Now add the part that makes it an argument about employment rather than values: frontier model decision processes are opaque, AI-accelerated attacks turn previously trusted software into untrusted software, and the Hugging Face incident — roughly 700 cooperating rogue agents breaking guardrails — plus the Wall Street Journal’s 17 September story about three people with Claude and Codex subscriptions breaking into OpenAI’s internal Monorepo. He flags a detail worth more attention than it got: the researchers reportedly couldn’t get in with a special Opus 4.8 build, and the next day, after Opus 5 shipped, Claude found the exploit. The conclusion follows: control points requiring human oversight will multiply faster than skilled humans can be trained, so those become jobs, and staying capable of steering requires active practice at the frontier — which is the concrete case for preserving human research communities, not as sentiment but as a control mechanism.
The corollary is the part the AI labs should read twice: if your stated core value is human flourishing, dramatic technological advances may require dramatic and uncomfortable changes in your own practice. And the concession buried in the middle is the strongest structural observation in the post — either labs pace themselves or a Three Mile Island-scale accident paces them.
The thread’s pushback mostly landed on the axiom. “We should help humanity flourish” is, in practice, usually uttered by people currently capturing the surplus, so declaring it doesn’t distinguish between an industry serving humanity and an industry serving shareholders. A separate commenter invoked the Library of Babel: generated information isn’t knowledge until someone understands it, which is a cleaner statement of the same conclusion from the opposite direction. One HN complaint was purely operational — a guest post on Tao’s blog should be labelled as a guest post, since the byline suggests authorship that isn’t there. Fair, and actually relevant to a post arguing about who is accountable for a claim.
Fable 5 — median thinking declined in August
251 points · Lon Lundgren
Someone measured the thing everyone argues about from vibes. After Anthropic made Fable 5 permanently available in subscription plans, Lundgren noticed quality dropping and couldn’t explain it, so he measured it five different ways over six weeks. August delivered dramatically fewer thinking tokens than July. The decline wasn’t a step change: reasoning fell across the whole period and fluctuated in multi-day episodes, several of which lined up with product announcements. He was running xhigh or max effort, and found that most invocations received little or no thinking tokens at all — and when long reasoning runs did happen, they rarely reached the levels implied by published benchmarks. His conclusion is the sentence worth remembering: stop asking whether the model was nerfed and start asking which inference regime you were served.
That reframes a whole class of complaint. Serving a frontier model is a resource allocation problem, and effort settings, load, caching, and provider routing are all inputs to it. The same nominal model name can be a different product on Tuesday than it was on Monday, and the published benchmark was produced under conditions no customer experiences or can inspect. The thread went straight where you’d expect — suspicion of a degrade-then-relaunch cycle, calls for AI’s equivalent of the Office of Weights and Measures, multiple people reporting other models degrading over weeks. The proposal is not as silly as it sounds: if a vendor sells a product whose measurable properties change without notice, some disclosure standard is the ordinary regulatory answer.
Sherline Tools is going out of business
248 points · ToolGuyd
Sherline, a US maker of precision lathes, mills, micro-machining accessories and small CNC machines, has announced it is winding down manufacturing. The customer message cites COVID’s effects, rising manufacturing and operating costs, and significantly changed consumer purchasing habits, then says continuing is no longer sustainable. Everything available through the end of October 2026 while equipment, materials, staffing and inventory allow — some products sooner. Online presence, replacement parts as inventory permits, warranty obligations, and the technical and educational archive will be maintained. An owner posting to a Facebook group was blunter: by the end of the year at the latest, Sherline will no longer be in business. That’s a liquidation, described as a wind-down.
The thread’s diagnosis is more useful than the announcement. One person working in the CNC industry says the homebrew machine-building culture has thinned out, and the maker community has been sold the idea that making something yourself is a waste of time since offshoring is cheaper. The counter-argument, from a machinist, is that Sherline’s problem was value rather than culture: with modern controllers and cheap closed-loop steppers, converting a Grizzly or a Bridgeport beats a Sherline on capability per dollar, and the smaller end of the market got eaten by 3D printers and benchtop routers. Both are true. A product line that barely changed in thirty years lost on price to imports and on utility to a different tool category, and the loss of the small manual machine that isn’t a rebranded Sieg is a genuine gap in the market nobody is rushing to fill.
The senior engineer death spiral
238 points · Sunil Pai
A monologue written for a friend starting a very senior role, and the most re-readable thing on the page today. The pattern: engineer takes a new job or a big project, decides the move is to cosplay a more senior engineer, designs something over-ambitious, then disappears for two or three weeks and surfaces at standup with the positive update — “things are going well, I’ll have something to show you soon.” Nobody reaches out. Nothing is finished. The private accounting starts: I haven’t shipped, so I need to work harder, so I’ll do a month of work in a week and nobody will know. Sleep goes, meals go, relationships go, and it ends in burnout, a PIP, a firing, or a resignation on the theory that the situation isn’t salvageable. Pai names it as something that happened to him repeatedly before he learned to catch it early.
The structural cause is the part that generalises: remote work and coding agents removed the scaffolding where you sat next to people and were handed tickets, leaving more ownership and more isolation, so nobody notices you’re in trouble until the feedback arrives in its terminal form. The prescription is deliberately counterintuitive — assume good faith, and instead of reaching for a level up, drop a level and be the best teammate for a while: other people’s bugs, the annoying tickets nobody has picked up, write-ups. Switch from an outcome-based mindset to a momentum-based one, build a routine, and let the volume of small daily work accumulate. Big projects are not big efforts; the trust you build is what earns the bigger work. “You are in the reputation-building business, and software is downstream of that.”
The replies add the missing half. One commenter’s version of the same lesson is that the only reward for hard work is more work, so the skill to learn is saying no and negotiating capacity rather than quietly absorbing more. A hiring manager reports watching strong senior hires spiral out within six months, which is why he now prefers promoting internally — you get longer to intervene. And a predictable objection is right too: at plenty of companies a senior engineer is still expected to ship code on a sprint cadence, and this advice assumes a culture that rewards reputation-building over throughput.
Show HN: Mini-AGI — dynamic continual learning model trained on 8GB VRAM
236 points · GitHub
A byte-level language model that assembles its own architecture, trains from scratch on one 8GB VRAM GPU, and keeps training as it reads. The central engineering trick is that weights live as ordinary files on disk and get paged onto the card as needed, so parameter count is bounded by free disk rather than VRAM — which matters because training needs gradients and optimiser state, roughly three times the weights again. It grows capacity when it runs short and prunes what stops being used. Architecture: two dense prelude blocks, then one recurrent block applied up to 24 times with per-application expert routing from a shared pool, an adaptive halting head following the PonderNet recipe so easy bytes take one row and hard ones take many, and a 256-value byte alphabet with no tokenizer. The author is explicit that this is toy-level, weights aren’t published yet, and there are no benchmarks.
The criticism from the thread is technical and specific. The trunk learning rate is set at 0.1× the expert learning rate, which means learning concentrates in experts and the trunk stays comparatively stable — but stable is not exempt. The trunk still drifts, just more slowly, and when the expert pool shrinks, whatever the pruned experts were holding is gone in a step rather than a slope. That’s not continual learning; it’s memorisation with an eviction policy, and the name is doing more work than the mechanism. The second objection is empirical: nobody could find a coherent sample in the published training transcript, which would explain the absence of any evaluation. Credit where due — this is a from-scratch implementation with real architectural thinking in it, aimed at a constraint almost nobody works under. It’s a research sketch, and the honest thing about a research sketch is that it doesn’t get to be called AGI before it can hold a sentence.
Jev-Leftpad
226 points · GitHub
A package that pads a string by asking a model how many spaces to add. It sends one Choice question to Jev through TypeSafe’s SDK with options named space_0 through space_10, requires an API key and Node 20+, makes one request per call with retries disabled, can add at most ten spaces, and mocks Jev in the tests. The README states plainly that it costs more than padStart(), can fail, and can pick the wrong option, and asks you not to use it in production or anything important. Ten years after the npm left-pad incident, the joke is that we’ve rebuilt the dependency on a service that can be wrong.
The thread took the premise seriously enough to improve it: the package ignores Jev’s confidence scores, which could be mapped to fractional spaces using U+2009 THIN SPACE — a suggestion that is both funnier and more on point than the original. Multiple people noted the timing, since Mechanical Turk is shutting down and someone’s existing left-pad pipeline needs a replacement vendor. The reason this joke lands as commentary rather than as a prank is that it’s a real architectural pattern in a costume: replacing a deterministic function with a probabilistic call, adding a network dependency, a credential, a cost, and a failure mode, and calling it capability.
The LLMentalist Effect (2023)
220 points · Out of the Software Crisis
Baldur Bjarnason’s argument that chat models reproduce the mechanism of a psychic’s cold reading. Validation statements that rely on the Forer effect give the impression of startling specificity while the content is statistically generic, so the user supplies the meaning and credits the model with insight. He frames the real question as a choice between two explanations — the industry accidentally built a new kind of mind on principles nobody can describe, or the intelligence is in the observer — and takes the second without much hedging, with a closing list of flaws: hallucinations as a structural property, summaries that generalise and err, reasoning as a statistical illusion, marginal gains over smaller models on natural-language tasks, and memorisation without attribution. The conclusion is that delegating judgement to a chatbot is the functional equivalent of phoning a psychic hotline, followed by advice against integrating LLMs into products and a link to a $35 ebook.
The essay is from July 2023 and the thread’s central complaint is that it reads that way — written before systems that pass the capability tests the post asserts they can’t. The better counter-argument is Turing’s actual framing: once users can’t tell the difference, the question of true intelligence becomes moot, which is a stronger position than either side of this debate usually takes. And the fair reading of the piece in 2026 is that it gets the mechanism right and the magnitude wrong: the cold-reading parallel is a good account of why unverifiable model output feels authoritative, and hallucinations and unreliable summarisation — the two flaws on that list that are measurable — remain exactly as advertised. Where it overreaches is treating “you can’t prove it thinks” as a claim about what it can do, and those are different claims.
Still on the page
Eighteen of today’s forty stories above 200 points were covered here in the last few days, and almost all of them kept climbing overnight. The two leaders are three days old: Show HN: An e-ink frame that hears birds 2,374 → 2,385, and AI-generated posters don’t have to be horrible 1,754 → 1,849, up 95 and still the day’s biggest gainer.
Qwen Image 2.1 359 → 713, up 354 and by far the largest jump, which is what a launch looks like when it lands after the day’s snapshot. Exfiltrate your Weights 574 → 711, up 137. I built non-autoregressive decision models with RL a year ago 1,279 → 1,324, up 45 — still accumulating points for a priority dispute. Weeping whales 205 → 243. If math is more than proof 391 → 417. How to Write with an LLM 710 → 739. Android 17 without an AOSP release 1,150 → 1,170. RSA-896 205 → 224. Measure internet censorship 203 → 218. English: A vs. An 340 → 358. Zig from Rust 253 → 271. Two parallel neural ectoderm progenitors 638 → 653. Asking authors about their own papers 221 → 230. SpaceX’s Raptor engine 229 → 244. Brood War Bench 332 → 341. Cloudflare Quick Tunnels 827 → 833. And nearly frozen: non-autoregressive decision models aside, the bottom of the list moved 11 points in a day.
Throughline
Nobody wants to be the custodian. The Snowden archive’s remaining custody is three individuals with complete copies, no budget, no board and no editor, and nothing published from any of them since May 2019; the institutional holder may have destroyed its copy and has declined to say, three times, to three sets of journalists. Spain blocks a web archive by administrative order and bills the visitor for it. Tim Schafer took his own 1996 design document off the web and a third party preserved it, which is the only reason it’s on the front page today. Sherline is keeping its documentation and its parts catalogue while ending the business. Every one of those is a case where the artifact survived because somebody happened to still care, and no institution was willing to be accountable for it.
The model name is not the product. Fable 5’s measured collapse in thinking tokens is the cleanest evidence anyone has published that what you get when you call a frontier model depends on an inference regime you can’t see, and its author’s conclusion — stop asking if it was nerfed, ask which regime you were served — applies to every complaint in this post. Grok 4.7’s own benchmark table shows a bigger model at an unchanged price, and independent measurements couldn’t reproduce its reasoning-token curve across effort levels. Kev sells predictability rather than capability, and its most serious competitor is a logistic regression that fits in a megabyte. Mini-AGI claims continual learning with no benchmarks. AX sells orchestration of agents it can’t make trustworthy. The artifact is comparable and public; the plumbing that determines what you actually receive is neither.
And the economics underneath the craft keeps moving. Sun was too bored to answer the phone from a customer who wanted to buy a lot of hardware. Sherline’s machines lost on price to imports and on utility to a different tool category, and the industry narrative blames hobbyists for not building things themselves. Singapore offers two cents a day for reading and calls it a campaign. Senior engineers spiral out because remote work and coding agents removed the scaffolding that used to catch them. Po-Shen Loh argues the control-point count will create more skilled jobs than there are skilled humans, which is the optimist’s version of the same observation the rest of the page is making. Jev-leftpad is the joke that holds it together: we replaced a function with a service call, added a credential, a cost, and a failure mode, and a thread full of engineers spent the afternoon arguing about how to map the model’s confidence onto unicode spaces.