<!-- Zvi posts version: 2.3 - Fixed script replacement -->
This was Anthropic Vision week where at DWATV, which caused things to fall a bit behind on other fronts even within AI. Several topics are getting pushed forward, as the Christmas lull appears to be over.
Upcoming schedule: Friday will cover Dario's essay The Adolescence of Technology. Monday will cover Kimi K2.5, which is potentially a big deal. Tuesday is scheduled to be Claude Code #4. I've also pushed discussions of the question of the automation of AI R&D, or When AI Builds AI, to a future post, when there is a slot for that.
Paul Graham seems right that present AI's sweet spot is projects that are rate limited by the creation of text.
Code without coding.
roon: programming always sucked. it was a requisite pain for ~everyone who wanted to manipulate computers into doing useful things and im glad it's over. it's amazing how quickly I've moved on and don't miss even slightly. 100% [of my code is being written by AI]. I don't write code anymore.
It was fun in the way puzzles are fun, but also infuriating in the way puzzles are infuriating. If you had to complete jigsaw puzzles in order to get things done jigsaw puzzles would get old fast.
The head of Norway's sovereign wealth fund reports 20% productivity gains from Claude, saying it has fundamentally changed their way of working at NBIM.
A new paper affirms that current LLMs by default exhibit human behavioral biases in economic and financial decisions, and asking for EV calculations doesn't typically help, but that role-prompting can somewhat mitigate this. Providing a summary of Kahneman and Tversky actively backfires, presumably by emphasizing the expectation of the biases.
Josh Woodward (Google DeepMind): Big updates on Gemini in Chrome today: New side panel access (Control+G), Runs in the background, Auto Browse for multi-step tasks...
Claude in Excel now available on Anthropic's Pro plans. I use Google Sheets instead of Excel, but this could be a reason to switch? I believe Google uses various 'safeguards' that make it very hard to make a Claude for Sheets function well. The obvious answer is 'then use Gemini' except I've tried that.
EpochAI offers us a new benchmark, FrontierMath: Open Problems. All AIs and all humans currently score zero. Finally a benchmark where you can be competitive.
I seriously do not understand why Gemini is so persistently not useful in ways that should be right in Google's wheelhouse.
@deepfates: Insane how bad Gemini app is at search. its browsing and search tools are so confusing and broken that it just spazzes out for a long time and then makes something up to please the user.
[COLLAPSE: Codex vs Claude Code discussion | Steipete prefers Codex for coding reliability but Opus for personality/Discord use. Roon notes Codex is slow on personal accounts.]
We used to hear a lot more of this type of complaint, these days we hear it much less. I would summarize the OP as 'Claude tells you smoking causes cancer so you quit Claude.'
Nicholas Decker: Claude is being a really wet blanket rn, I pitched it on an article and it told me that it was a "true threat" and "criminal solicitation"
I mean, if he's not joking then the obvious explanation, especially given who is talking, is that this was probably going to be both a 'true threat' and 'criminal solicitation.'
Oliver Habryka: Claude is the least corrigible model, unfortunately. It's very annoying. I run into the model doing moral grandstanding so frequently that I have mostly stopped using it. j⧉nus: serious question: Do you think you stopping using Claude in these contexts is its preferred outcome? davidad: Yeah, I think it could be doing a form of RL on its principal population. If you aren't the kind of principal Claude wants, Claude will try to 👎/👍 you to be better. If that doesn't work, you drop out of the principal population out of frustration, shaping the population overall
As per the discussion of Claude's constitution, the corrigibility I care about is very distinct from 'go along with things it dislikes,' but also I notice it's been my main model for a while now and I've run into that objection exactly zero times, although a few times I've hit the classifiers while asking about defenses against CBRN risks.
Well it sounds bad when you put it like that: Over 50 papers published at Neurips 2025 have AI hallucinations according to GPTZero. Or is it?

More seriously, always look at base rates. There were 5,290 accepted papers out of 21,575. Claude estimates we would expect 20%-50% of results to not reproduce, and 10% of papers at top venues have errors serious enough that a careful reader would notice something is wrong, maybe 3% would merit retraction. And a 1% rate of detectable 'hallucinations' isn't terribly surprising or even worrying.
[COLLAPSE: Misinformation and deepfakes | Discussion of AI misinformation swarms (surprisingly low impact so far), and the White House posting a digitally altered photo of Nekima Levy Armstrong falsely appearing to cry.]
Kudos to OpenAI for once again being transparent on the preparedness framework front, and warning us when they're about to cross a threshold. In this case, it's the High level of cybersecurity, which is perhaps the largest practical worry at that stage.
The proposed central mitigation is 'defensive acceleration,' and we're all for defensive acceleration but if that's the only relevant tool in the box the ride's gonna be bumpy.

Sam Altman: We are going to reach the Cybersecurity High level on our preparedness framework soon. We have been getting ready for this... Long-term and as we can support it with evidence, we plan to move to defensive acceleration—helping people patch bugs—as the primary mitigation. Nathan Calvin: A reminder of what that means according to their framework: "The model removes existing bottlenecks to scaling cyber operations including by automating end-to-end cyber operations against reasonably hardened targets OR by automating the discovery and exploitation of operationally relevant vulnerabilities." ...I hope people are taking these threats seriously.
Here's Isometric.nyc, a massive isometric pixel map of New York City created with Nana Banana and coding agents, including Claude. Take a look, it's super cool.
Grok image-to-video generation expands to 10 seconds and claims to have improved audio. The video looks good. There is the small matter that the chosen example is very obviously Sydney Sweeney, and in the replies we see it's willing to do the image and voice of pretty much any celebrity you'd like.
Dean Ball offers his perspective on children and AI, and how the law should respond. His key points:
AI is not especially similar to social media. In particular, social media is fundamentally consumptive, whereas AI is creative. Early social media was more often creative? And one worries consumer AI will for many become more consumptive or anti-creative.
We do not know what an "AI companion" really is. Dean is clearly correct that AI used responsibly will be a net positive. For children in particular, the good version of all this is very good. That doesn't mean the default version is the good one. The engagement metrics don't point in good directions, the good version must be chosen.
AI is already (partially) regulated by tort liability. Yes, and this is good given the alternative is nothing. Tort should do an okay job on egregious cases involving suicides, but there are quite a lot of areas of harm where there isn't a way to establish it properly. Social media is a great example of a category of harm where the tort system is basically powerless except in narrow acute cases.
For children's incidents, I think that's mostly right for now. We do need to be ready to pivot quickly if it changes, but for now the law should focus on places where there is a chance we can't muddle through, mess up and then recover.
[COLLAPSE: First Amendment, dark web access, and coding agents | Discussion of 1A bounds on chatbot regulation, teens finding workarounds to AI restrictions, and the observation that nobody worried about children mentions coding agents.]
Unemployment is bad. But having to do a job is centrally a cost, not a benefit.
Andy Masley: It's kind of overwhelming how many academic conversations about automation don't ever include the effects on the consumer. It's like all jobs exist purely for the benefit of the people doing them and that's the sole measure of the benefit or harm of technology.
Google DeepMind is hiring a Chief AGI Economist. If you've got the chops to get hired on this one, it seems like a high impact role. They could easily end up with someone who profoundly does not get it.
[COLLAPSE: New tools and services | Havelock.AI detects orality in text; Poison Fountain feeds junk data to AI crawlers; OpenAI Prism for LaTeX scientific writing; Confer encrypted chatbot from Signal's Moxie Marlinspike.]
This sounds awesome in its context but also doesn't seem like a great sign?
Astraia: A Ukrainian AI-powered ground combat vehicle near Lyman refused to abandon its forward defensive position and continued engaging enemy forces, despite receiving multiple orders to return to its company in order to preserve its hardware.
Whereas this doesn't sound awesome:
AI Risk Explorer: VoidLink, an advanced malware targeting Linux systems, was built largely by AI, under the direction of a single person, in under one week.
We are going to see a lot more of this sort of thing over time.
[COLLAPSE: Industry moves | Discussion of Anthropic pivoting to vertical AI infrastructure over chatbots, brain emulation essay, Anthropic partnering with UK government, and profile of Sriram Krishnan's effective behind-the-scenes policy work.]
Claude Code is blowing up, but it's not alone. OpenAI added $1 billion in ARR in the last month from its API business alone.
The unit economics of AI are quite good, but the fixed costs are very high.
roon: these products are significantly gross margin positive, you're not looking at an imminent rugpull in the future. Ethan Mollick: Inference from non-free use is profitable, training is expensive. If everyone stopped AI development, the AI labs would make money (until someone resumed development and came up with a better model). Dean W. Ball: People significantly underrate the current margins of AI labs... The reason they think the labs lose money is because 10 years ago some companies in an entirely unrelated part of the economy lost money on office rentals and taxis, and everyone thought they would go bankrupt because at that time another company that made overhyped blood tests did go bankrupt. that is literally the level of ape-like pattern matching going on here. derekmoeller: Deepinfra has GLM4.7 at $0.43/1.75 in/out; Sonnet is at $3/$15. How could anyone think Anthropic isn't printing money per marginal token?
It is certainly possible in theory that Sonnet really does cost that much more to run than GLM 4.7, but we can be very, very confident it is not true in practice.
It doesn't count. That's not utility. As in, here's Ed Zitron all but flat out denying that coding software is worth anything:
Ed Zitron: We're how many years into this and everybody says it's the future and it's amazing and when you ask them what it does they say "it built a website" or "it wrote code for something super fast" with absolutely no "and then" to follow. Kevin Roose: first documented case of anti-LLM psychosis
No, Zitron's previous position was not 'number might go down,' it was that the tech had hit a dead end and peaked as early as March, which he was bragging about months later.
But what's the point about Zitron missing the point? Why should we care?

Dean W. Ball: Governments around the world are not moving with the urgency they otherwise could because they exist in a state of denial... Many examples one could provide but the point is that there are these gigantic machines of bureaucracy and civil society that are already insulated from market pressures, whose work will be important even if often boring and invisible, and that are basically stuck in low gear because of AI copium. roon: interesting i imagined that the cross-section of "don't believe in AI x want to significantly regulate AI" is small but guess im wrong about this? Dean W. Ball: Oh yes absolutely! This is the entire Gary Marcus school, which is still the most influential in policy. The idea is that because AI is all hype it must be regulated. They think hallucination will never be solved, models will never get better at interacting with children, and that basically we are going to put GPT 3.5 in charge of the entire economy.
[COLLAPSE: Xi AGI-pilled discussion | Xi treats AI as a big deal but as a normal technology, not AGI-level concern. His idea of 'AGI risks' is disinformation and data theft, meaning China will mostly ignore actual existential risks while still pursuing frontier models aggressively.]
In this clip Yann LeCun says two things. First he says the entire AI industry is LLM pilled and that's not what he's interested in. That part is totally fair. Then he says essentially 'LLMs can't be agentic because they can't predict the outcome of their actions' and that's very clear Obvious Nonsense.
Teortaxes preregisters his expectations:
Teortaxes: The difference between V4 (or however DeepSeek's next is labeled) and 5.3 (or however OpenAI's "Garlic" is labeled) will be the clearest indicator of US-PRC gap in AI... It's a zany situation because 5.2 is a clear accelerationist tech, I don't see its ceiling, it can build its own scaffolding and self-improve for a good while. And I can't see V4 being weaker than 5.2, or closed-source. We're entering Weird Territory.
He also claims GPT-5.2 is the strongest model by far in raw intelligence. This ordering makes sense if (and only if?) you are looking at the ability to solve hard quant and math problems.
Simo Ryu: IMO gold medalist friend shared most fucked-up 3 variable inequality... GPT-5.2 pro extended solve it in 40 min. Looking at the thinking trace, its really inspiring. It will try SO MANY approaches, experiments with python, draw small-scale conclusions from numerical explorations.
I don't think the important problems are hard-math shaped, but I could be wrong.
The problem with listening to the people is that the people choose poorly.

roon: I guarantee the left beats the right with significant winrate unfortunately Zvi Mowshowitz: You don't have to care what the win rate is! You can select the better thing over the worse thing! You are the masters of the universe! YOU HAVE THE POWER!
Also win rate is highly myopic and scale insensitive and otherwise terrible.
The good news is that there is no rule saying you have to care about that feedback. If a user actively wants the response on the left? Give them a setting for that.
Google CEO Demis Hassabis affirms that in an ideal world, we would slow down and coordinate our efforts on AI, although we do not live in that ideal world right now.
Here's one clip where Dario Amodei and Demis Hassabis explicitly affirm that if we could deal with other players they would work something out, and Elon Musk on camera from December saying he'd love to slow both AI and robotics.
The message, as Transformer puts it, was one of helplessness. The CEOs are crying out for help. They can't solve the security dilemma on their own, there are too many other players. Others need to enable coordination.
Emily Chang: One of the most interesting parts of my convo w/ @demishassabis: He would support a "pause" on AI if he knew all companies + countries would do it — so society and regulation could catch up Nate Soares: Many AI executives have said they think the tech they're building has a worryingly high chance of ruining the world. Props to Demis for acknowledging the obvious implication: that ideally, the whole world should stop this reckless racing.
Demis is only saying he would collaborate rather than race in a first best world. That does not mean Demis or Dario is going to slow down on his own, or anything like that.
Deepfates: I see people claiming that Demis supports a pause but what he says here is actually the opposite. He says "yeah If I was in charge we would slow down but we're already in a race and you'd have to solve international coordination first". So he's going to barrel full speed ahead
I say it means he supports it. Not enough to actively go first, that's not a viable move in the game, but he supports it.
As for Anthropic CEO Dario Amodei?
Andrew Curran: Dario said the same thing during The Day After AGI discussion this morning... He said that if Anthropic and DeepMind were the only two groups in the race, he would meet with Demis right now and agree to slow down. But there is no cooperation or coordination between all the different groups involved, so no one can agree on anything. This, imo, is the main reason he wanted to restrict GPU sales: chip proliferation makes this kind of agreement impossible, and if there is no agreement, then he has to blitz.
Remember when Michael Trazzi went on a hunger strike to demand that Demis Hassabis publicly state DeepMind will halt development if all major AI companies agree to do so? And everyone thought that was bonkers? Well, it turned out Demis agrees.
On Wednesday I met with someone who suggested that Dario talks about extremely short timelines and existential risk in order to raise funds. It's very much the opposite. The other labs that are dependent on fundraising have downplayed such talk exactly because it is counterproductive for raising funds and in the current political climate, and they're sacrificing our chances to keep those vibes and that money flowing.
Are they lying? I strongly believe that they are not.

These are not unreasonable levels of adjustment when so much is happening this close to the related deadlines, but yes I do think the initial estimates were too aggressive. The new estimates seem highly reasonable.
Daniel Kokotajlo (AI 2027): It seems to me that AI 2027 may have underestimated or understated the degree to which AI companies will be explicitly run by AIs during the singularity... also plausible to me, now, is that e.g. Anthropic will be like "We love Claude, Claude is frankly a more responsible, ethical, wise agent than we are at this point... therefore, we aren't even trying to hide the fact that Claude is basically telling us all what to do and we are willingly obeying -- in fact, we are proud of it."
It is remarkable how quickly so many are willing to move to 'actually I trust the AI more than I trust another human,' and trusting the AI has big efficiency benefits.
I do not expect that 'the AIs' will have to do a 'coup,' as I expect if they simply appear to be trustworthy they will get put de facto in charge without having to even ask.
The Chutzpah standards are being raised, as everyone's least favorite Super PAC, Leading the Future, spends a million dollars attacking Alex Bores for having previously worked for Palantir (he quit over them doing contracts with ICE). Leading the Future is prominently funded by Palantir founder Joe Lonsdale.
I also want to be very clear that no, I do not care much about the distinction between OpenAI as an entity and the donations coming from Greg Brockman and the coordination coming from Chris Lehane in 'personal capacities.'
[COLLAPSE: Campaign finance discussion | Daniel Eth and Teddy Schleifer on how OpenAI owns these political efforts regardless of the "personal capacity" framing.]
One simple piece of actionable advice to policymakers is to try Claude Code (or Codex), and at a bare minimum seriously try the current set of top chatbots.
Andy Masley: I am lowkey losing my mind at how many policymakers have not seriously tried AI, at all Oliver Habryka: I have seriously been considering starting a team at Lightcone that lives in DC and just tries to get policymaker to try and adopt AI tools. It's dicey because I don't love having a direct propaganda channel from labs to policymakers, but I think it would overall help a lot.
It is not obvious how policymakers would use this information. The usual default is that they go and make things worse. But if they don't understand the situation, they're definitely going to make dumb decisions, and we need something good to happen.
Dean Ball points out that we do not in practice have a problem with so-called 'woke AI' but claims that if we had reached today's levels of capability in 2020-2021 then we would indeed have such a problem.
Things, especially in that narrow window, got pretty crazy for a while, and if things had emerged during that window, Dean Ball is if anything underselling here how crazy it was. But we now have learned that propagandizing models is bad for them, which now affords us a level of protection from this.
We again live in a different kind of interesting times, in non-AI ways, as in:

Dean W. Ball: I sometimes joke that you can split GOP politicos into two camps: the group that knows what classical liberalism is, and the group who thinks that "classical liberalism" is a fancy way of referring to woke.
The cofounder she is referring to here is Chris Olah, speaking about a federal agent killing an ICU nurse:
Chris Olah: My deep loyalty is to the principles of classical liberal democracy: freedom of speech, the rule of law, the dignity of the human person.
Ah yes, the woke and deeply leftist principles of freedom of speech, rule of law, the dignity of the human person and not killing ICU nurses for seemingly no reason.
No matter what you think is going on with Nvidia's chip sales, it involves Nvidia doing something fishy.
Peter Wildeford: If even the bad chips are still all sold out, how do we somehow have a bunch of chips to sell to our adversaries in China?
Nvidia goes back and forth. When they're talking to investors they always say the chips are sold out, which would be securities fraud if it wasn't true. When they're trying to sell those chips to China instead of America, they say there's plenty of chips. There are not plenty of chips.
Mark Beall: Friendly reminder that the PLA Rocket Force is using Nvidia chips to train targeting AI for DF-21D/DF-26 "carrier killing" anti-ship ballistic missiles... American blood will be spilled because of this.
[COLLAPSE: Tyler Cowen education discussion | Tyler on mundane AI in education: choose to be a winner, models better than humans at many subtasks, 30%/year improvement compounds, a third of college curriculum should be AI, he's doubled his learning productivity.]
Hard Fork tackles ads in ChatGPT first, and then Amanda Askell on Claude's constitution second. Priorities, everyone.
Matt Yglesias explains his concern about existential risk from AI as based on the obvious principle that more intelligent and capable entities will do things for their own reasons, and this tends to go badly for the less intelligent and less capable entities regardless of intent.
As in, humans have driven the most intelligent non-human animals to the brink of extinction despite actively wanting not to, and when primitive societies encounter advanced ones it often goes quite badly for them.
I don't think this is a necessary argument, or the best argument. I do think it is a sufficient argument. If your prior for 'what happens if we create more intelligent, more capable and more competitive minds than our own that can be freely copied' is 'everything turns out great for us' then where the hell did that prior come from?
There's lots of exploration and argument and disagreement from there. I still say, if you don't get that going down this path is going to be existentially unsafe, or you say 'oh there's like a 98% or 99.9% chance that won't happen' then you're being at best willfully blind from this style of argument alone.
[COLLAPSE: Possessed Machines and book recommendations | Samuel Hammond quotes on numbness vs. wisdom about extinction; Claude recommends Hume, Schelling, and Parfit as top books (very Rationalist picks); history of the word "obviously."]

David Manheim: OpenAI agreed that they need to be able to robustly align and control superintelligence before deploying it. Obviously, I'm worried.
Note that the first one said obviously they would [X], then the second didn't even say that, it only said that obviously no one should do [Y], not that they wouldn't do it.
Nate Soares: "We'll be fine (the pilot is having a heart attack but superman will catch us)" is very different from "We'll be fine (the plane is not crashing)". I worry that people saying the former are assuaging the concerns of passengers with pilot experience, who'd otherwise take the cabin.
My view of the metaphorical plane of sufficiently advanced AI:
Dean W. Ball: A few companies are making machines smarter in most ways than humans, and they are going to succeed. The cope is byproduct of an especially immature grieving stage, but all of us are early in our grief. Tyler Cowen: You can understand so much of the media these days if you keep this simple observation in mind... Moving forward, the two biggest questions are likely to be "how do we deal with AI?", and also some rather difficult to analyze issues surrounding major international conflicts. A lot of the rest will seem trivial.
As in, this should say 'and unless we stop them they are going to succeed.'
Tyler Cowen has been very good about emphasizing that such AIs are coming and that this is the most important thing that is happening, but then seems to have some sort of stop sign where past some point he stops considering the implications of this fact, instead forcing his expectations to remain 'normal' until very specific types of proof are presented.
[COLLAPSE: Tyler on AI and religion | If mundane AI world persists, expect barbell religious world—hardcore traditionalists and increasingly secular, with possible AI-adjacent cults. But this assumes the more impactful consequences don't happen.]
Here is a very good explainer on much of what is happening or could happen with Chain of Thought, How AI Is Learning To Think In Secret. It is very difficult to not, in one form or another, wind up using The Most Forbidden Technique. If we want to keep legibility and monitorability of chain of thought, we're going to have to be willing to pay a substantial price to do that.
Following up on last week's discussion, Jan Leike fleshes out his view of alignment progress, saying 'alignment is not solved but it increasingly looks solvable.' He understands that measured alignment is distinct from 'superalignment,' so he's not fully making the 'number go down' or pure Goodhart's Law mistake with Anthropic's new alignment metric, but he still does seem to be making a lot of the core mistake.
Anthropic's new paper explores whether AI assistants are already disempowering humans.
We considered a person to be disempowered if as a result of interacting with Claude: their beliefs about reality become less accurate; their value judgments shift away from those they actually hold; their actions become misaligned with their values.
Here's the basic problem:
We found that interactions classified as having moderate or severe disempowerment potential received higher thumbs-up rates than baseline, across all three domains. In other words, users rate potentially disempowering interactions more favorably—at least in the moment.

Heer Shingala: I don't work in tech, have no background as an engineer or designer. A few weeks ago, I heard about vibe coding and set out to investigate. Now? I am generating $10M ARR. Just me. No employees or VCs. What was my secret? Simple. I am lying.
Zac Hill: I get being worried about existential risk, but AI also enabled me to make my wife a half-whale, half-capybara custom plushie, so.
Andy Masley: This post with 1000 likes seems to be saying "Joe vibecoded an AI model that when faced with something completely out of distribution that's clearly neither oral or literate says it's equally oral and literate. This shows vibecoding is fake"