Hacker Newsnew | past | comments | ask | show | jobs | submit | bestcommentslogin
Most-upvoted comments of the last 48 hours. You can change the number of hours like this: bestcomments?h=24.

(I work at Anthropic)

Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.

Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.

[1] https://github.com/harbor-framework/terminal-bench-science


> Meanwhile, the over thinkers on Hacker News come up with convoluted reasons to hate on Firefox every time the subject arises...Firefox is our last best hope for browser engine diversity and competition.

It's exactly because Firefox is so important that you'll see people here complaining when Firefox does things to push away users like Mozilla buying up an ad-tech company, collecting data on users, and using firefox to push personalized ads, or the addition of anti-features and questionable design choices that force us to hunt for and modify poorly documented settings in about:config and make edits to userChrome.css

It isn't bots complaining about Firefox here, it's power users who are frustrated by what Firefox is turning into. Users who are seriously concerned about what Mozilla is prioritizing, and who are genuinely worried about what is at stake.

I hope people here never stop bitching about Firefox. Refusing to talk about Firefox's problems wont help make them go away. Keep discussing what you'd like to see in Firefox and what things you hope they'll focus on and prioritize. There are Firefox devs and mozilla employees around here. If we're lucky, a few of them might see and listen to some of what we say. It's the most tech savvy users who disable all the telemetry and data collection, so our feedback isn't really going to be seen any other way.


One thing I’m observing in these comments is a willingness of folks to project their own predictions onto Ed’s statements when validating their plausibility. Eg. “I think he’s wrong about the timing but I do expect AI companies to go to zero.”

You can do that, but then you’re no longer discussing his predictions. You’re discussing your predictions, and your own positioning.

Those differ from Dan’s essay, which engages with the literal text of Ed’s numerous predictions during 2024 and 2025 which are demonstrably invalidated by their measurable outcomes.


I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle.

Aced it, got the job as a senior software engineer.

The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".


The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.

Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":

https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f

Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...


Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI.

I feel the same about Grok w/ Elon. I will pay extra to use someone else.

I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.

And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.


Hi! Author here. Surprised to see this on HN now. Happy to answer any questions!

Some context about this:

- This is NOT an LLM. its a small ar transformer trained from scratch. One of the points was that extremely complex problems can be tackled without LLMs

- Till the v1 of this result, this benchmark was only scaled by LLMs or their finetunes (ofc w enormous training costs). Other attempts performed okayish but used v complex architectures or extremely high amounts of training compute. No one expected a simple AR transformer to perform this well, at this low cost and w these few training samples.

- Sample Efficiency is one of the most important unsolved problems today in AI. That's what I was targetting with this work. We know it is easy to increase SE by increasing compute/params, so it was important to constrain cost as much as possible (also why OpenAI's Parameter Golf had fixed compute and why Modded NanoGPT is considered very sample efficient)

- Can the perf be improved? Yes but the competition is ongoing so can't talk about it

- Personally I think today's frontier models can be beat by training from scratch. Haven't proved this yet tho

- Fun: I was new to ML when I posted this first (dec '25). I basically used ARC as a way to learn ML


LWN is one of the, if not the single, highest signal tech publications around. I hope they're able to maintain a stable subscription service. Being user funded, and avoiding having to maintain allegiance to advertisers, is likely part of why their quality is so high.

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.”

I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better.

I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.


The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).

This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.

Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:

Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.

GDPval-AA v2: +1.5% vs Opus 5.

OSWorld 2.0: +2.5% vs Opus 5.

Humanity's Last Exam (with tools): +1.6%

Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?


Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization. This even caused some of our (testing) prompts to be used for training (which Mistral removed after we expressed our disappointment).

For some time these pages conflicted with what our users reported (they said that in contrast to what I stated to our management they found they were opted into training on prompts by default as per their own privacy page). Mistral just now corrected their docs. I'm not sure how long the conflicting situation has lasted, but at least for several days.

For contrast: Claude disables training on prompts for organizations starting from the 18 euro tier [0]. As a European I'm disappointed.

[0] https://claude.com/pricing#team-&-enterprise


It has always seemed to me that they're hacking for dopamine response in moderately interested data labelers.

> Postal Service mail ballot system has catastrophic problems

Yes. That's the point. The administration doesn't care about fixing whatever hypothetical mail ballot fraud may or may not exist; they are doing this to inject chaos into the election results, giving them cover to claim any result they want or even to just throw out the results entirely. Catastrophic problems are the means to that end.


Apps need to exit the Play Store. I have a whole slew of apps which haven't been updated for a while now because Aurora Store isn't working because of whatever Google of doing. Meanwhile the ones installed via FDroid and other means are perfectly up to date. It's a crappy state of affairs as Google continues to close the boiling Android frog.

There’s an old mantra from organizing: “no permanent enemies, no permanent allies.” You will never align with another group or person 100% on all things; the secret to making change is to build a coalition of those with whom you agree on an issue without holding it against them that you disagree on a different one.

I disagree with Mozilla about many things, but I agree with this article - I use Firefox because it’s the only browser out there that isn’t Chrome or WebKit, and that’s worth enough to me that I’m willing to disagree with them on other issues.


I would be interested in seeing a similar list of predictions from Altman, Amodei, etc with annotations about how many have come true. Ed Zitron is a blow hard and frequently overstates things to the point where it is hard to take seriously, but so are the AI industry leaders.

I have heard multiple breathless press releases warning that the end of white collar work is "just 6 months away" and that people not using the latest Mythos/Fable/Whatever model will be hopelessly left behind.


I am under the impression that Firefox is the only web browser which has access to quality ad blocker. Am I incorrect in this? How is this not enough of a selling point for everyone to switch to it?

Slightly tangential thought:

My belief is: legislation needs to either make it just as hard to merge two companies as it is to unmerge them, or make it just as easy to unmerge two companies as it is to merge them.

It's insane to me that for how often companies merge and cause competition issues, we effectively never see the opposite happen. I know there's a ceremonial approval for merging two companies (at least in the US), but it's just impossible to undo or prevent the damage.


Would be nice if they donate to LibreOffice then, to improve the support of various MS Office features in files, as well as comparison/diffing features. Win-win to everyone.

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

I'm still waiting for effort max to finish.

EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!

Excerpts from the reasoning trace:

> Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.

> Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.

> I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]

> I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]

> Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...]

> I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean.

This is a notable result because most of the recent Claude models have been pretty bad at drawing pelicans, at least when compared to models in the Gemini or GLM series.


If I recall correctly, there were some papers which suggested that LLMs favor LLM-generated passages over human written ones. I can consistently reproduce this by asking Claude which code snippet it prefers: the one it generated in a different chat, or one that I refactored for my own needs and find more useful. It always picks its own :) I've also experienced that both Claude and Codex routinely include generated websites when I ask them to search for something. It also doesn't help that the web search tools that OAI and Anthropic have are deeply limiting: can't exclude keywords or domains.

I find them almost unintelligible. I'm a native English speaker. I read a lot, so I think my comprehension should be at least OK. I'm not even particularly stupid. Yet when faced with things like below (a direct copy/paste from a handoff document in a long running vibe-coding session), I have no real idea of what it's trying to tell me. Is it important? Do I need to do anything?

I think that spending all day trying to parse stuff like this is why a long session is so exhausting

> Worth stating because four documents now assert it. The console freeze was recorded in exactly one place with exactly one justification — a dead drag handle during a booked half-day you do not get back — and handoff-4.3-done.html's own wording is that 4.4's review page "could not break the console, but the downside of being wrong is that half day". No second reason. Checked, not recalled.


> They're packing lots of signal into fewer words

There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are writing for each other then presumably they could stick to CoT-speak (unless it's a distillation risk?).


Zitron has become the distorted reflection of the very AI boosters he criticizes and mocks.

I think the worst thing that happened to him was AI skepticism becoming a political position. This gave him a captive audience - as long as he says what they want to hear, which means that he can never ever concede that he might have been wrong or that AI might actually be progressing or having successes.

This is not conducive to good prediction long-term - rather it leads one to a state of cognitive dissonance where one's chosen enemies must be simultaneously terrifyingly powerful and incompetent dunces. The propagandist's disease.


Sometime around 1980-81 I had a part time job while an undergraduate in college doing system programming/admin for the Caltech High Energy Physics department.

Rob Pike was the system programmer/admin before me when he was a grad student in high energy physics, but he left to go work at Bell Labs.

One day another student, Karl Heuer, and I both were engaging in the common programmer pastime of complaining about the screen editors of the day and saying we could write something better.

Somehow this turned into a competition, and we both spent all night racing against each other writing our editors. It was mostly silent except for the typing, interrupted by the occasional announcement of some feature that was now working to hopefully rattle the other.

In the morning the other student system programmer/admin, Norman Wilson, got in and saw what Karl and I had been up to.

Norman mentioned this in an email to Rob Pike. His response was something close to this:

> Everyone writes a screen editor. It's easy to do and makes them feel important. Tell them to work on something useful.

It was only a couple years or so later that Rob Pike wrote a screen editor. I wonder if it made him feel important? :-)


I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks.

I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since left ceased bothering with human languages, and it takes a rare sort of direct matrix-gazing savant to be able to try and horse-whisper them into doing or revealing anything they didn't already plan to do.


Zhuangzi was fishing in the Pu River when the King of Chu sent two high officials to him. They said, “The king wishes to entrust you with the affairs of his realm.”

Zhuangzi kept hold of his fishing rod and did not turn around. He said, “I have heard that in Chu there is a sacred turtle that has been dead for three thousand years. The king keeps it wrapped in cloth, enclosed in a box, and stored in the ancestral temple.

“Now, would this turtle rather have died and left its bones behind to be honored, or would it rather be alive, dragging its tail through the mud?”

The two officials replied, “It would rather be alive, dragging its tail through the mud.”

Zhuangzi said, “Then go away. I too will drag my tail through the mud.”


Definitely cool.

I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.


I didn't want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high):

  llm logs -cx | llm -m claude-fable-5.1 -s 'animate this'
Here's the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

It's excellent!


When do we stop this instinctive response of "well you can still do it in an only slightly convoluted way" everytime a corporate does something bad.

Windows added ads - well you can disable them, if you don't like it.

Chrome brought up mv3 - well you can still use mv2, it is only optional.

Reddit locking subreddits behind login wall - well there's always old.reddit

Android moving everything to playstore - there's always FDroid.

How many of these are still true and for how long?

If someone's country is a dictatorship, it doesn't help to tell them that there are 100 other democratic countries, just move. Not everyone can emigrate or want to emigrate for any number of reasons.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: