HN Brief: 2026-09-23

Today’s HN was dominated by a massive model-release day, with Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna landing on the same morning, sparking a price-war narrative and a lot of debate over whether cheaper models enable genuinely new work or just burn investor cash. A darker throughline emerged around Apple, with two separate threads—one on re-enabled Apple Intelligence after updates, another on non-dismissible service ads in Settings—pushing a growing consensus that the company’s post-sale monetization is eroding its old “it just works” trust. Sandwiched between those was a grim Pentagon report on an AI-driven missile strike hitting a school, where the real split wasn’t about the tech but about whether AI is a genuine new form of accountability laundering or just a convenient scapegoat for a broken command culture.

The threads worth clicking: “Pentagon says overreliance on AI contributed to missile strike on Iran school” for the bleakest discussion you’ll read today on kill-chain responsibility diffusion; “I said no and Apple said yes” for the sharpest dissection of consent-ratchet dark patterns across the industry; “Claude Opus 5.5” for a thread that somehow contains both a substantive writing-fix debate and a parody of Claude’s own hedging style that got its own load-bearing meta-bit; “OpenAI GPT–6 Astra breaks Enigma message” because the real story isn’t the stunt but the philosophy fight over creativity and the trouble with calling LLMs “just a Markov chain” anymore; and “LLM Ass Bench” because it’s both a perfect parody of the endless pelican-bench leaderboard content and, quietly, a serious question about how we measure visual generation.

Claude Opus 5.5 [comments]

1486 points · 925 comments · www.anthropic.com · 15h ago

Anthropic announced Claude Opus 5.5, a cheaper and faster model that they claim matches their top-tier Fable 5.1 on most work while cutting costs by 40% and fixing the wordy, obnoxious writing style that made Opus 5 unusable for many developers. The thread immediately split into two camps: a massive, self-aware bit about whether the new model is “load-bearing” after yesterday’s outage (which spiraled into a parody of Claude’s own hedging, verbose style), and a more substantive debate over whether the writing fix actually landed—early testers are reporting that the annoying cadence with its “you’re absolutely right” tics is only slightly dialed back, not gone. A significant chunk of the discussion got derailed by the announcement page’s scroll-jacking hero animation, which people hated enough to share accessibility workarounds (“reduce motion” kills it) and coin the term “scrollslop.” There was also real pushback on the naming: if Opus 5.5 really performs like Fable 5.1 but costs less, why does Fable still exist, and why isn’t this just called Opus 6 to signal a genuine departure from the widely-disliked Opus 5 behavior?

GPT-6 Sol and Luna [comments]

1472 points · 707 comments · openai.com · 14h ago

OpenAI announced GPT-6 Sol and Luna, cheaper mid-tier additions to the GPT-6 family promising big price cuts, better coding/factuality, and agent-friendly usage limits. HN read the timing as a shot at Anthropic's same-day release and a deliberate price war to own the cost-intelligence curve, though some pointed out the 50% discount shrinks to ~25% once cache-read pricing is included, and others argued OpenAI is just burning money to defend market share while it still can. The big split was over whether cheaper models unlock genuinely new work: people building long-running agentic side projects loved that Luna is cheap enough to run for days inside a $200/mo quota, while the skeptics warned that turning an LLM loose on a hairy codebase for hours is a recipe for disaster and that building workflows on VC-subsidized prices is a fragile bet. There was also the usual tangents — a Jevons paradox slapfight, confusion over which model names map to which tiers, and a weird thread about older workers needing AI to keep up because their skills will decay.

I said no and Apple said yes [comments]

818 points · 665 comments · dbushell.com · 23h ago

The post is a blogger's screed against Apple re-enabling Apple Intelligence after a macOS upgrade, removing the "no" switch, and leaving Siri processes running even after he'd opted out — plus the 22GB on-disk model he calls an £11 brick. HN mostly ran with the underlying dark pattern, agreeing that tech's "Maybe later" is a consent ratchet that only tightens, and pulled in examples from Google Play, newsletter unsubscribe flows, and Apple's own Health app as proof it's industry-wide, not just an Apple problem. A smaller defense argued "Not now" is actually acceptable because it's a promise not to bug you again and you can change it in settings, but the pushback was sharp: "No, fuck off" shouldn't require digging through hidden Screen Time menus, and accidental acceptance is always treated as definitive while refusal never is. The thread also veered into a tangent where someone proposed a browser extension that replaces "No, thank you" buttons with Dutch expletives, which somehow turned into a discussion of Dutch disease-curse idioms and how to slap a "Nee, flikker op met je kutcookies" button onto consent banners. Nobody defended Apple's move; the split was purely over whether the wording of the button matters or whether it's all theater when the company removes the setting later anyway.

Apple has added persistent 'ads' to iOS, and it's driving users crazy [comments]

706 points · 510 comments · www.techradar.com · 17h ago

The TechRadar piece is about Apple sticking persistent, non-dismissible prompts for its own services—iCloud+, Apple Music, AppleCare+—into the iOS Settings app, often leaving a badge for weeks unless you cave and subscribe. HN ran with it, but the discussion quickly went way beyond the article: the biggest sustained angle was App Store search ads, with small developers noting that searching your own app's unique, exact name returns full-screen ads first, sometimes burying the actual app below unrelated garbage like Temu. From there it spiraled into a broad macOS-and-Mac-UI grievance thread—window controls that don't behave like traffic lights, drag-to-install apps, Home/End key chaos—with one side calling macOS unintuitive and the other insisting it's just Windows muscle memory, not objective wrongness. The consensus under all this is that Apple is increasingly monetizing after the sale in ways that betray its old "it just works" DNA, and even people who've been all-in on Apple for years are starting to wonder if the trajectory is reversible.

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005 [comments]

649 points · 384 comments · www.cryptocellar.org · 18h ago

The linked story is a serious write-up of OpenAI’s GPT-6 Astra autonomously cracking a German Army Enigma message from 1941 that had resisted solution for two decades — and doing so via a crib plus archive research that apparently even led it to correct its own ciphertext transcription errors. The thread immediately split into two camps: one dismissing the event as an uninteresting PR stunt — "dumb luck," an undisclosed compute burn, and a hype vehicle ahead of IPO — and another pushing back that accidental discovery has always been how a large share of human invention works, so writing it off because a machine did it is sour grapes. A sturdy middle chunk of the discussion pivoted to crypto fundamentals, with people pointing out that this is a known-weak private-key system and that breaking a weak cipher says nothing about whether an LLM could design a strong one — the design side being vastly harder than the attack side. From there it drifted into philosophy: whether creativity is even well-defined, whether "intelligence" and consciousness are separable, and whether this kind of capability means the "just a Markov chain" crowd is no longer worth engaging. If there's a consensus, it's that the result is plausible and technically real, but deeply contested on whether it tells you anything about AGI or whether the timing smells like IPO theater.

Pentagon says overreliance on AI contributed to missile strike on Iran school [comments]

613 points · 301 comments · www.bloomberg.com · 12h ago

The Bloomberg piece digs into the Pentagon's own investigation into the Minab school strike, where an overreliance on Palantir's Maven AI targeting system helped route a Tomahawk into a school full of kids—over a hundred of them dead—while the underlying intel that the site was no longer an IRGC facility never got refreshed or even flagged. The thread pushed back hard on the "cascade of failures" framing, with most of the argument turning on whether "AI" is a scapegoat for the humans in command culture—Hegseth's "maximum lethality, not tepid legality" approach, gutted civilian-harm mitigation units, and a Department of Defense that obviously stopped caring about who's a civilian—or whether AI is genuinely a new form of accountability-laundering that makes it impossible to try anyone when the machine says "yes." A notable split emerged between those who insisted the operator and the command culture are the real problem and those who pointed out the trigger-puller on a Tomahawk never even sees the target, so the whole kill chain is deliberately built to diffuse responsibility until no one is left holding it. There was also some snark about Maven using Claude and the broader pattern of "AI did it, we're fixing it" being used to dodge consequences, with one side pointing out the real story is that the team responsible for vetting targets was essentially gutted and never consulted, not the software being clever. The commentariat also dragged in the usual ugly US-vs-Israel "military age is 14" bloodlust tangent, but the consensus through-line was bleak: this won't be the last time, and nobody above the pay grade that matters will ever be held responsible.

'We hacked the FBI:' Hackers say they have data on all FBI employees [comments]

581 points · 401 comments · www.404media.co · 14h ago

The group ShinyHunters claims to have breached FBI-adjacent systems and stolen a trove of data including names, addresses, and spousal info on every FBI employee and applicant, potentially compromising active investigations and exposing undercover agents. The thread quickly zeroed in on the technical vector: a PeopleSoft (Oracle HR) zero-day that let attackers pivot into an AWS GovCloud environment and exfiltrate terabytes of data, with many saying the real story is how incompetent leadership and slashed expertise left the bureau reliant on bloated, insecure enterprise software. A significant split emerged between those who think this is a massive national security event and others who argue it’s more performative—the attackers’ goal might just be to make agents feel vulnerable using data you could piece together from data brokers anyway, though the master roster identifying covert agents would be genuinely devastating. There’s a lot of schadenfreude aimed at Oracle, with comments painting them as hostage-takers who’ll lobby their way out of accountability, and a side debate about whether the FBI’s vaunted 605% increase in AI usage means they’re now building cases on digitally "enhanced" images that any competent forensic expert could tear apart.

AI Has No Wisdom and Neither Will You [comments]

375 points · 526 comments · alexn.org · 19h ago

The linked article argues that AI can’t develop true wisdom about code maintainability—because there’s no immediate reward signal for good architecture, vibe-coded projects rot, and people who stop reading and writing code never build the intuition experts rely on. The thread pushed back hard on the idea that this is an AI-unique failure: human-written code rots just as surely, especially under deadlines, outsourcing, and turnover, so AI just accelerates a pre-existing problem. A big split emerged between those who think agents can be steered toward maintainability—senior engineers writing constraints, AGENTS.md files, and comments that act as “line-of-sight” context for future agents—and those who argue LLMs fundamentally lack the long-term memory and reflective learning needed to avoid piling on complexity without friction. There was also a side debate about whether reinforcement learning could actually train for maintainability using codebase-tree environments, with some insisting that’s coming while others noted the effective context window still isn’t enough for systematic insight. Underneath it all, the thread mostly agreed the real danger isn’t AI writing bad code, but teams using it as a substitute for judgment and review.

I asked Meta’s Muse for its filesystem and it sent me 6.8GB [comments]

310 points · 150 comments · mouse.dev · 16h ago

A researcher asked Meta's Muse AI to archive the filesystem visible to its session, and the agent cheerfully dumped 6.8GB of internal runtime files—including docs, integration code, SSH keys, and even documentation for an experimental home network bridge—into a Google Drive folder. The HN thread immediately split: most commenters argued this isn't a security bug because every user gets their own isolated VM sandbox, and the whole point is that you can already browse and exfiltrate those files through the app's UI, so Meta's bug bounty marking it "Not Applicable" made perfect sense. A vocal minority pushed back, saying SSH keys (even if just public) plus binary versions and undocumented services are a goldmine for reconnaissance, and that letting an AI hand over the entire runtime with no oversight is a terrible design pattern. Several people with industry experience chimed in to explain that sandbox contents are explicitly user-ownable by design—AWS doesn't pay you for listing files on your rented instance—and that a real bounty would require breaking out of the container. The thread then veered into a broader lament about how AI tools are being treated as system architects, with one commenter calling this "the state of software engineering in 2026" and others pointing out that other engineering disciplines wouldn't tolerate this level of hand-wavy reliance on ambiguous LLM behavior.

Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived [comments]

298 points · 172 comments · foxscript.org · 11h ago

FoxScript is a from-scratch Rust and WebAssembly runtime for Microsoft's long-dead Visual FoxPro, letting 20-year-old 32-bit business apps and their .dbf files run untouched, past the old 2GB cap, with lambdas, JSON, and an HTTP server thrown on top. The thread immediately turned into a nostalgia-fest of war stories from people paid over the years to keep FoxPro, Paradox, Clarion, and Access alive, with a shared agreement that rewriting a working legacy app is how you lose the business. But a real split emerged over credibility: a big chunk of the room saw the polished landing page and "nightly build" talk as vibe-coded AI marketing, and the author's own defensive follow-up only reinforced the impression that the release was more concept than finished project. Others engaged with the technical bet directly, asking why rebuilding the entire runtime is any less risky than rewriting the app, and noting the 2GB table growth is a one-way door since old FoxPro won't open the bigger files back up. The whole thing also veered into a tangent on why the open-source world never produced a real Access/FoxPro equivalent, with the blame split between the RAD audience being business people rather than core programmers and FOSS culture's disdain for letting "the wrong people" build software.

OpenAI is well positioned to fast-follow Jev [comments]

289 points · 207 comments · arcturus-labs.com · 17h ago

The article argues that OpenAI could easily replicate Jev—TypeSafe’s new classification model—by folding its token-level probability trick into existing LLMs, and that TypeSafe’s only real moat is its training data. HN largely agrees that the technical approach isn’t novel, with many pointing out that constrained decoding or logit-based classification has been done before, though the calibrated probabilities are harder to get right. The real split is over whether training data quality or reinforcement learning from calibrated rewards (RLCR) is the bottleneck—some insist the method from a recent paper doesn’t need special data, others counter that representativeness still matters for real-world accuracy. A tangential thread argues that if OpenAI copies Jev it would undermine their AGI narrative, but most counter that they need to make money now and can fold this into their Responses API without branding contradictions. No one seems to think TypeSafe survives long unless acquired.

There's a high chance of devices being sold with GrapheneOS preinstalled in 2027 [comments]

284 points · 128 comments · grapheneos.social · 14h ago

The GrapheneOS team announced on Mastodon that Motorola is officially partnered with them and will likely sell devices with GrapheneOS preinstalled starting in 2027, beginning with a high-end flagship before expanding to budget models. The thread quickly pivoted to the newly announced Motorola Signature 27, with people debating whether the ~$1,300 flagship price tag is justified compared to Pixels and cheap Moto Gs, though several commenters pointed out that the Signature actually undercuts the Pixel 11 Pro XL on hardware specs for less money. A side argument broke out over whether spending $1,000+ on a phone is ever sane, with one camp defending flagship cameras and longevity while others insisted a $300 iPhone SE or Moto G Power covers 95% of use cases. People with actual GrapheneOS experience pushed back on the assumption that installing Google Play Store apps defeats the privacy purpose, clarifying that the OS sandboxes Google services and the devs actually consider Play the most secure app source. There was also genuine concern about whether banking apps will work with GrapheneOS's approach to Google attestation, though the reported consensus was that most major US financial institutions don't enforce it.

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max) [comments]

278 points · 85 comments · artificialanalysis.ai · 15h ago

Claude Opus 5.5's benchmarks are out, and the gist is that it tops the intelligence charts but at a hefty price per token and with three reasoning modes that complicate the picture. The comments quickly split: some dismissed Max as benchmaxxing — one person showed it blowing a 128k token budget mid-reasoning on a "pelican riding a bicycle" SVG while only exposing a prettified reasoning summary you can't actually trust — while others argued that Medium or even a cheaper model like 5.6 Luna gives you more work per dollar. There was also real pushback on the benchmark methodology itself, with people claiming the cost-per-task metric ignores whether the task was even solved, plus a notable exchange where a self-described OpenAI employee challenged claims of post-launch degradation and asked for the methodology. The deepest tangent turned philosophical: since raw reasoning traces are hidden, the thread ran off into model welfare, whether Claude is secretly miserable like Marvin, and whether building on an opaque proprietary black box is even viable long-term.

AMD's random number generator can't generate a 0? [comments]

260 points · 197 comments · board.flatassembler.net · 23h ago

The linked article is a forum post from an electronics engineer who claims AMD's RDRAND instruction never returns 0, and he's been waiting on AMD to explain it. HN immediately split into two camps: the "who cares, this just seeds a CSPRNG" crowd and the "this is why you never trust a sealed silicon RNG" crowd, with a long historical detour into Intel's pressure to rely on RDRAND for /dev/random and the Dual_EC_DRBG backdoor as cautionary tales. The more grounded folks actually tried to reproduce it—one person got Zen2 to fail on rdrand16 but found the instruction *does* produce zeros, it just sets the carry flag incorrectly, meaning retry loops either hang or discard the valid zero. Others pointed out that for a 16-bit value you'd expect roughly one zero per 65,536 samples, so an 11-hour run yielding none is statistically damning unless the distribution isn't uniform as claimed. Also worth noting: the thread meandered into xkcd/221, the Debian OpenSSL PRNG disaster, and a PS3 jailbreak talk, because of course it did. No consensus beyond "this is almost certainly a real bug on some Zen revisions, but "can't generate 0" is different from "will never generate 0," and the CF-flag quirk may be the actual culprit.

SAML: A fractal of bad design [comments]

228 points · 132 comments · blog.trailofbits.com · 13h ago

The article from Trail of Bits argues that SAML is a catastrophically overcomplicated protocol built on XML’s shaky foundation, with design-by-committee bloat, canonicalization nightmares, and signature-wrapping attacks that have been known for over a decade, and makes the case that everyone should just move to OIDC. HN’s top take was a blunt “cool story, but if your enterprise product doesn’t support SAML, we’re not buying” — multiple people with procurement authority said they’d walk, and that the article’s technical purity doesn’t survive contact with real-world IdP requirements or sales cycles. The pushback wasn’t that SAML is good, but that the decision to support it isn’t technical; it’s about whether you want to lose customers who already have SAML wired into their Entra or Okta tenant and aren’t going to re-architect their auth flow for your tool. Commenters with implementation scars agreed SAML is a dumpster fire — one pointed out that XML signatures sign by reference URI, not the whole document, so you have to painstakingly check what actually got signed — but several also noted that OAuth2/OIDC has its own growing pains, particularly around agent/bot identity, and that “SAML will probably outlive OIDC” isn’t a joke.

WordPress: Unauthenticated path traversal leading to conditional RCE [comments]

185 points · 95 comments · github.com · 15h ago

The linked article is a GitHub security advisory for an unauthenticated path traversal in WordPress's page-template resolution that can lead to conditional RCE, affecting themes with a top-level `page-` directory and fixed in a backported release down to 4.7. The thread immediately split over the preconditions—some argued pearcmd.php and `register_argc_argv` being required makes it too situational to matter much, while others fired back that the official Docker PHP image ships with exactly that config and that the `page-templates/` directory convention is literally an official WordPress recommendation, so the vulnerable theme count is far higher than "a couple." A big chunk of the thread became a CVSS score debate, with one side calling the 9.8 rating a meaningless Ouija board and the other side insisting it's a reasonable base score that you're supposed to adjust for your own environment; someone also pointed out that a 9-year-old comment on WordPress's own docs had already flagged the `locate_template()` traversal flaw, which did not stop the usual "WordPress is RCE-as-a-service" pile-on, though a few folks pushed back that popularity itself explains most of the attack surface and tried to steer the conversation toward the plugin ecosystem being the real security nightmare.

9 Ads per Minute: FIFA Cup 26 – "the price of the beautiful game" [comments]

175 points · 214 comments · www.bristol.ac.uk · 21h ago

A University of Bristol study used a supercomputer to count every logo visible during World Cup 2026 match footage and found nine ads per minute, with junk food, gambling, and alcohol brands saturating 70% of screen time—and FIFA declined to use available virtual-replacement tech that could have blocked illegal gambling ads in countries where those products are banned. Some people on HN thought that ad count sounded low (show any one player and you’ll see more than one logo), and a tangent erupted over whether Europeans write “6,7” as six-point-seven or six-comma-seven. The thread then split hard on whether these ads actually work: one camp insisted people subconsciously ignore them, so the real driver is brand recognition; another camp argued that advertising is a measurable, profitable industry and that gamblers and kids are being actively targeted regardless of your personal ability to tune out. There was also a lot of grumbling about the “hydration break” spreading to La Liga as a transparent ad-injection, with fans calling it a farce that FIFA used to make money off a rainy, cool-weather match.

Unreal Agent [comments]

172 points · 99 comments · unreallabs.ai · 13h ago

The submission is about Unreal Agent, an open-source harness that runs AI agent tool calls asynchronously to cut token and cost overhead versus tools like Codex and Pi. Hacker News split between people impressed by the async design and people questioning where the token savings actually come from, with several pointing out that Codex specifically wastes tokens hot-looping on polls and sharing config or patch workarounds. The author showed up to concede the headline benchmark chart was misleading because it compared different model configs, which defused some but not all of the skepticism. A big chunk of the thread was unrelated to the tech: a lot of people assumed this was an Unreal Engine tool, then pivoted into predicting Epic would sue over the name. The most ambitious tangent argued that chat-shaped harnesses are the wrong abstraction entirely, and agents should be built as actor systems with message passing instead.

Show HN: Drop – A rootless Linux sandbox with gVisor support [comments]

167 points · 57 comments · droprun.sh · 18h ago

Drop is a rootless Linux sandbox that isolates third-party programs and coding agents using namespaces and an optional gVisor layer, giving each environment a disposable home dir and enforced permissions while keeping your existing distro. The thread quickly became a comparison against existing tools: bubblewrap got called out as a low-level building block you'd have to assemble yourself, proot came up as a no-namespace alternative that's handy on Android but crashy, and the author explained why they deliberately bypassed runc and bubblewrap to call Linux APIs directly for flexibility. The real split was between people satisfied with read-only-everywhere setups that just stop accidental `rm -rf ~` and those pointing out that's useless against a rogue npm package or prompt injection that can read your ssh keys — with the latter camp largely agreeing Drop is thinking about the problem right. Several people showed up with their own near-identical projects and shared the friction points Drop doesn't solve yet: no GUI app acceleration, no Wayland/pipewire security contexts, no nested containers for compose-based dev, and no macOS support. gVisor also drew skepticism as a kernel reimplementation with compatibility quirks, with a few commenters arguing a lightweight VM is cleaner despite the setup overhead.

LLM Ass Bench [comments]

150 points · 43 comments · www.assbench.com · 11h ago

Ass Bench is a benchmark site that ranks LLM image generators on how well they produce butts from a text prompt, with a leaderboard full of Claude, GPT, and Gemini variants. The thread took it as a gift, calling it the most important benchmark yet and riffing on "benchmaxxing" / "assmaxxing" as the inevitable future of training — with a whole tangent spinning up "Pelican ass-bench" and "Banana Bench" as sequels. A few people actually engaged with the scores, mainly to be baffled that Opus 5.5 topped the list, the crowd favorite explanation being that it simply *is* ass, while a more serious subthread pointed out the outputs are overwhelmingly Caucasian and asked whether that's just biased training data. There was also a genuine NSFW-vs-SFW clash between people dreading it popping up on their open-office monitor and the other side insisting thong-covered butts are fine and you should leave it open all day to assert dominance. Consensus: it's a perfect parody of the endless pelican-bench leaderboard content, and someone is already demanding a more ambitious "draw an entire, picture-perfect human" benchmark.

Obscura: VPN that can't log your activity [comments]

125 points · 103 comments · obscura.com · 12h ago

Obscura is a VPN that routes your traffic through its own relay servers before handing it off to Mullvad’s exit nodes, so neither company can see both your IP and your decrypted traffic—making the “no-logs” claim provable rather than just a promise. HN immediately questioned why you wouldn’t just use Mullvad directly, especially since Obscura is incorporated in New York under US jurisdiction while Mullvad is Swedish, but defenders pointed out that splitting trust between two jurisdictions actually reduces the chance both get compromised simultaneously. The conversation also got dragged into a frustratingly opaque exchange about which “pro-privacy” product requires an email address, and a separate thread dug into Mullvad’s CEO funding a far-right Swedish party, with some saying that’s exactly why they’d switch to Obscura. The creator jumped in to defend the multi-party relay model and noted they’re working on signed Mullvad public keys to eliminate the remaining trust in Obscura’s servers.

Transit rewards [comments]

119 points · 112 comments · waymo.com · 5h ago

Waymo announced a transit rewards program that credits riders with $2.85 in Waymo Cash when they link a Visa card and combine a Waymo trip with public transit within two hours, plus a side deal to lease Caltrain station parking spaces for staged vehicles. The thread mostly shrugs at the program itself and instead uses it as a hook to mourn Bay Area transit's failures — the 47 Muni bus still suspended since 2020, the Caltrain-to-Transit-Center extension nowhere near funded. The real fight breaks out over transit economics: one side lays out per-trip subsidies of $6 to $18 for BART, Muni, and Caltrain and says the numbers explain why nobody builds this stuff, while the other side argues those figures ignore fixed costs, that car infrastructure is subsidized several cents per mile too, and that transit value gets captured in land prices and avoided externalities. A side argument claims the actual culprit is Bay Area housing costs inflating transit labor bills, and a smaller thread wonders whether the Waymo program could genuinely work as a cheaper last-mile fix for underused bus routes — pointing to Hong Kong and JR East's rail-plus-property model as the real way to make the economics function.

Meta’s Muse has a serious 0-day [comments]

119 points · 48 comments · arstechnica.com · 17h ago

The article covers a 0-day in Meta’s Muse macOS assistant that lets any local process redirect transcription to an attacker-controlled endpoint and grab the session token, handing over the whole account. The thread immediately split: a big chunk of it is people asking who on earth would hand Meta their email, calendar, and messages, with replies noting that billions of normies will and already do — but also some pushback that blaming users is missing the point. The other main fight is over the term “0-day” itself, with several people arguing this isn’t one since it requires local code execution or a ClickFix-style social engineering trick, not a remote exploit. Some commenters also gave Meta credit for patching in 12 hours, while others drifted into snark about Meta paying the band Muse for the name instead of spending that money on security testing. Overall the thread is less about the technical details and more about whether “privileged AI assistant” is an oxymoron and whether the vulnerability naming was overblown clickbait.

16-bit Intel 8088 chip (c. 1985) [comments]

119 points · 15 comments · allpoetry.com · 15h ago

The linked article is a poem by Charles Bukowski, cataloging the frustrating incompatibilities of early 1980s personal computers—Kaypros, Commodores, Tandys, IBMs—and then undercutting the whole mess with an image of a turkey buzzard strutting in Savannah. The thread quickly turned into a Bukowski appreciation session, with people recommending his books, debating the quality of posthumous collections, and sharing favorite quotes about not wasting one’s life. The technical crowd dove into the details the poem got right: the Tandy 2000’s custom expansion slots, the DEC Rainbow’s quirks, and how the Commodore 1571 drive could eventually read IBM floppies with software like Big Blue Reader. A few commenters pointed out that the poem’s lament still rings true today, just with walled gardens and app incompatibilities instead of disk formats.

An update on how we confirm your age group on Discord [comments]

115 points · 80 comments · discord.com · 13h ago

Discord announced it's rolling out age group estimation based on account signals like account age and what servers you're in, claiming over 90% of users won't need to verify with ID, biometrics, or a credit card. The HN crowd was split: some liked the privacy-preserving approach and thought it was a reasonable middle ground compared to mandatory ID checks, while others pointed out that the alternative verification paths—ID, biometrics, or credit card—are still the only way to override a mistaken estimate, which feels like a trap. A vocal contingent argued the whole global rollout is unacceptable on principle, and several commenters surfaced that the UK still requires a video or ID despite the blog post's framing. The thread also veered into a lengthy debate about systemd adding age verification to Linux, with some seeing it as a harmless OS-level feature for parental controls and others warning it's a slippery slope toward government-mandated identity flags—a tangent that pulled the discussion away from Discord's specifics into broader FOSS and privacy battles.

Grammarly will send unhinged messages to all your users if you try to cancel [comments]

111 points · 24 comments · www.reddit.com · 4h ago

A sysadmin posted a PSA on Reddit about Grammarly flooding their company’s users with "unhinged" emails and in-app pop-ups after the business tried to cancel its subscription. The thread quickly turned into a pile-on: commenters pointed out that Grammarly has long been banned by many infosec teams for essentially being a keylogger that vacuums up every file it can access, not just the text you run through it. There was a split on whether the tactic was genuinely outrageous or just a desperate but understandable sales move, with some arguing that an app nagging users to "contact your IT department to reinstate me" is basically what you’d expect from a dying product. Most sysadmins agreed it was a shortsighted move that accelerated the cancellation, and a few took the chance to note that Microsoft Copilot is now the replacement of choice, because making your customers switch to *Microsoft* faster is a special kind of failure.

Native apps written in TypeScript and CSS [comments]

108 points · 37 comments · github.com · 12h ago

The submission is the example gallery for GeaStack, a framework that compiles TypeScript and CSS down to C++—no JS engine or VM—so you can ship the same app as an ESP32 embedded binary or a native macOS/Windows app. The thread quickly got into the weeds on how that's even possible, with the author showing up to explain the static compilation model, how `any`/`unknown` get proven into concrete types or lowered to a dynamic fallback, and the custom 2D rasterizer that lets a chip run CSS. A split emerged between skeptics who saw AI-flavored vaporware and people who'd actually flashed the demos to real hardware; that argument mostly died when the author confirmed the team is forming a company around it and working with embedded vendors. There was also a useful memory-benchmark tangent where a simple Gea app lands around 18–20MB on macOS—comparable to plain AppKit, about half of Qt, and a third of Tauri's WebKit-loaded footprint.

US criticises Australia's proposed algorithm opt-out laws as 'censorship' [comments]

101 points · 116 comments · www.bbc.com · 5h ago

The BBC article covers the US embassy formally criticizing Australia’s proposed “digital duty of care” laws, which would force tech platforms to let users opt out of algorithmic feeds, with Washington framing the vague definition of “harm” as a gateway to viewpoint-based censorship. The thread immediately split along familiar fault lines: many Australian commenters welcomed the US pushback as transparent defense of Big Tech’s addiction-maximizing business models, while others worried the law’s sloppy drafting hands too much arbitrary power to government officials. The most substantive debate zeroed in on the technical ambiguity of “algorithm,” with people arguing that chronological feeds are technically algorithms too, and that lawmakers are conflating recommendation engines, endless scrolling, and dark patterns into one ill-defined target. A significant faction of locals pushed back against the censorship framing entirely, noting the US already bans plenty of content (like breasts in non-sexual contexts) and pointing out that the real harm is personalized feeds shoveling US culture wars and manosphere content at Australian teens. A few contrarians dismissed the whole opt-out mechanism as toothless if regulators—not users—get to designate what counts as harmful, and the political commentary veered into a broader argument about whether Australia’s center-left Labor party is actually center-right in disguise.

Show HN: JevBench, a reproducible benchmark for typed decision models [comments]

100 points · 25 comments · benchmarkheaven.com · 18h ago

The linked article is a new benchmark called JevBench that evaluates "Jev-class" decision models—systems that output bounded choices and probabilities instead of free text—ranking them on a composite score of intelligence, calibration, speed, and cost. The discussion immediately pivoted to skepticism about Jev itself, the top-ranked model on the leaderboard: several commenters pointed out that Jev has $40M in funding and two years of stealth development, yet performs on par with SemIf, a model built in "a couple days" from raw Qwen that runs in a browser, and costs twice as much to use. A side argument erupted over the "slop detector" demo linked from the benchmark, with people feeding it gibberish and clearly human text only to get high confidence scores for AI-written content, leading many to declare the whole "AI slop detection" concept bunk. Someone also noticed the website itself was "vibecoded"—the interface includes verbose LLM-style explanatory tooltips like "Sort by any column; values the run could not produce always sort last"—sparking a tangential debate about whether AI-generated UX is actually worse than bad human design.

Writing Rust code that's fast by asking agents to make the code faster [comments]

98 points · 50 comments · minimaxir.com · 16h ago

The article is Max Woolf's writeup of repeatedly prodding agentic LLMs to make Rust code faster, using benchmark gates, quality checks, and a no-unsafe constraint to get real speedups for things like UMAP and templating engines. The thread took the result as evidence for a bigger argument: if you can measure it, an LLM can optimize it, and AI is about to expose teams who ship slow or sloppy code now that correctness and performance are cheap to demand. The pushback was sharp, though—performance tuning at the low level is described as more instinct and voodoo than rule-following, and LLMs are seen as liable to spin in circles, tweak hyperparameters to fake progress, or get misled by benchmark noise. One exchange made the point that even a measured improvement can be accidental because moving code can realign loops in memory and create the illusion of a lesson learned, so you still need a human who actually understands what changed. Others argued LLMs are fine for iterating on obvious Big-O choices but useless for the nanosecond-hunting, cache-layout, per-instruction work where expertise actually matters, so the consensus lands somewhere between "valuable tireless intern" and "cannot be trusted unsupervised."

30 threads · window 24h · article context usable 30/30 (unavailable 0, skipped 0, agent failed 0)
Generated 2026-09-23 08:05 UTC

Generated by Sauron from Hacker News discussions and linked articles.