<!-- Zvi posts version: 2.3 - Fixed script replacement -->
Anthropic released a new constitution for Claude. I encourage those interested to read the document, either in whole or in part. I intend to cover it on its own soon.
There was also actual talk about coordinating on a conditional pause or slowdown from CEO Demis Hassabis, which I also plan to cover later.
Claude Code continues to be the talk of the town, the weekly report on that is here.
OpenAI responded by planning ads for the cheap and free versions of ChatGPT.
There was also a fun but meaningful incident involving ChatGPT Self Portraits.
Tone editor or tone police is a great AI job. Turn your impolite 'f you' email into a polite 'f you' email, and get practice stripping your emotions out of other potentially fraught interactions, lest your actual personality get in the way. Or translate your neurodivergent actual information into socially acceptable extra words.
If your query is aggressively pattern matched into a basin where facts don't matter and you're making broad claims without much justifying them, AIs will largely respond to the pattern match, as Claude did in the linked example. And if you browbeat such AIs about it, and they cower to tell you what you want to hear, you can interpret that as 'the AI is lying to me' or you can wonder why it decided to do all of that.
Claude adds four new health integrations in beta: Apple Health (iOS), Health Connect (Android), HealthEx, and Function Health. They are private by design.
OpenAI adds the ChatGPT Go option more broadly, at $8/month. If you are using ChatGPT in heavy rotation, you need to be paying at least the $20/month for Plus to avoid being mostly stuck with Instant.
The pitch is that Gemini now draws insights from across your Google apps to provide customized responses.
The pitch is that it can gather information from your photos (down to things like where you travel, what kind of tires you need for your car), from your Email and Google searches and YouTube and Docs and Sheets and Calendar, and learn all kinds of things about you, not only particular details but also your knowledge level and your preferences. Then it can customize everything on that basis.
One potential 'killer app' is fact finding. If you want to know something about yourself and your life, and Google knows it, hopefully Gemini can now tell you.
The real killer app would be taking action on your behalf. It can't do that except for Calendar, but it can do things on the level of writing draft emails and making proposed changes in Docs.
When such things work, they 'feel like magic.'
When they don't work, they feel really stupid.
I asked for reactions and got essentially nothing.
That checks. To use this, you have to use Gemini. Who uses Gemini?
Thus, in order to test personalized intelligence, I need a use case where I need its capabilities enough to use Gemini, as opposed to going back to building my army of skills and connectors and MCPs in Claude Code, including with the Google suite.
Olivia Moore: Connectors into G Suite work just OK in ChatGPT + Claude - they're slow and can struggle to find things. If Gemini can offer best "context" from Gmail, G Drive, Calendar - that's huge.
The other problem is that Google's connectors to its own products have consistently, when I have tried them, failed to work on anything but basic tasks. Even on those basic tasks, the connector from Claude or ChatGPT has worked better. And now I'm hooking Claude Code up to the API.
Elon Musk and xAI continue to downplay the whole 'Grok created a bunch of sexualized deepfakes in public on demand and for a time likely most of the world's AI CSAM' as if it is no big deal. Many countries and people don't see it that way, investigations continue and it doesn't look like the issue is going to go away.
A Bloomberg report says 'one in eight kids personally knows someone who has been the target of a deepfake video.' Reports rose from roughly 4,700 in 2023 to over 440,000 in the first half of 2025.
We could stop Grok if we wanted to, but the open-source tools are already plenty good enough to generate sexualized deepfakes and will only get easier to access. You can make access annoying and shut down distribution, but you can't shut the thing down on the production side.
Meanwhile, psychiatrist Sarah Gundle issues the latest warning that this 'interactive pornography,' in addition to the harms to the person depicted, also harms the person creating or consuming it, as it disincentivizes human connection by making alternatives too easy, and people (mostly men) don't have the push to establish emotional connections. I am skeptical of such warnings and concerns, they always are of a form that could prove far too much and historical records mostly don't back it up, but on the other hand, don't date robots.
Misinformation is demand driven, an ongoing series.

Jerry Dunleavy IV: Neera Tanden believes that ICE agents chased a protester dressed in Viking gear and sitting in a bath tub with skateboard wheels down the street, and that the Air Force was called in in response. Certain segments of the population just are not equipped to handle obvious AI slop.
This is not a subtle case. The chyron is literally floating up and down in the video. In a sane world this would be a good joke. Alas, there are those on all sides who don't care that something like this is utterly obvious, but it makes little difference that this was an AI video instead of something else.
In Neera's defense, the headlines this week include 'President sends letter to European leaders demanding Greenland because Norway wouldn't award him the Nobel Peace Prize.' Is that more or less insane than the police unsuccessfully chasing a bathtub viking on the news while the chyron slowly bounces?
The new OpenAI image generation can't do Studio Ghibli properly, but as per Roon you can still use the old one by going here.
Roon: confirmed that this is a technical regression in latest image model nothing has changed WRT policy.

Sienna Rose recently had three songs in the Spotify top 50, while being an AI, and we have another sighting in Sweden.
The technical name for this edition is 'ads in ChatGPT.' They attempt to reassure us that they will not force sufficiently paying customers into the Nexus, and it won't torture the non-paying customers all that much after all.
Sam Altman: We are starting to test ads in ChatGPT free and Go (new $8/month option) tiers. Here are our principles. Most importantly, we will not accept money to influence the answer ChatGPT gives you, and we keep your conversations private from advertisers. It is clear to us that a lot of people want to use a lot of AI and don't want to pay, so we are hopeful a business model like this can work.



So, on the principles:
If you wanted to know what 'AGI benefits humanity' meant, well, it means 'pursue AGI by selling ads to fund it.' That's the mission. I do appreciate that they are not sharing conversations directly with advertisers, and the wise user can clear their ad data. But on the free tier, we all know almost no one is ever going to mess with any settings, so if the default is 'share everything about the user with advertisers' then that's what most users get. Ads not influencing answers directly, and not optimizing for time spent on ChatGPT are great, but even if they hold to both the incentives cannot be undone.
Also, we saw the whole GPT-4o debacle, we have all seen you optimize for the thumbs up. Do not claim you do not maximize for engagement. And you know Fidji Simo is itching to do it all.
This was inevitable. It remains a sad day, and a sharp contrast with alternatives.
Then there's the obvious joke:
Alex Tabarrok: This is the strongest piece of evidence yet that AI isn't going to take all our jobs.
I will point out that actually this is not evidence that AI will fail to take our jobs. OpenAI would do this in worlds where AI won't take our jobs, and would also do this in worlds where AI will take our jobs. OpenAI is planning on losing more money than anyone has ever lost before it turns profitable. Showing OpenAI is not too principled or virtuous to sell ads will likely help its valuation, and thus its access to capital, and the actual ad revenue doesn't hurt.
As you would expect, Ben Thompson is taking a victory lap and saying 'obviously,' also arguing for a different ad model.
Ben Thompson: The advertising that OpenAI has announced is not affiliate marketing; it is, however, narrow in its inventory potential (because OpenAI needs inventory that matches the current chat context) and gives the appearance of a conflict of interest (even if it doesn't exist). What the company needs to get to is an advertising model that draws on the vast knowledge it gains of users — both via chats and also via partnerships across the ecosystem that OpenAI needs to build — to show users ads that are compelling not because they are linked to the current discussion but because ChatGPT understands you better than anyone else.
I think Ben is wrong. Ads, if they do exist, should depend on the user's history but also on the current context. When one uses ChatGPT one knows what one wants to think about, so to provide value and spark interest you want to mostly match that. Yes, there is also room for 'generic ad that matches the user in general' but I would strive as much as possible for ads that match context.
What, Google sell ads in their products? Why they would never:
Demis Hassabis told me Google has no plans to put ads in Gemini. "It's interesting they've gone for that so early," he said of OpenAI putting ads in ChatGPT.
roon: big fan of course but this is a bit rich coming from the research arm of the world's largest ad monopoly
Parmy Olson calls ads 'Sam Altman's last resort,' which would be unfair except that Sam Altman called ads exactly this in October 2024.
Starting out your career at this time and need a Game Plan for AI? One is offered here by Sneha Revanur of Encode. Your choices in this plan are Tactician playing for the short term, Anchor to find an area that will remain human-first, or Shaper to try and make things go well. I note that in the long term I don't have much faith in the Anchor strategy, even in non-transformed worlds, because of all the people that will flood into the anchors as other jobs are lost. I also wouldn't have faith in people's 'repugnance' scores on various jobs:

People can say all they like that it would be repugnant to have a robot cut their hair, or they'd choose a human who did it worse and costs more. I do not believe them. What objections do remain will mostly practical, such as with athletes. When people say 'morally repugnant' they mostly mean 'I don't trust the AI to do the job,' which includes observing that the job might include 'literally be a human.'
Anthropic's Tristan Hume discusses ongoing efforts to create an engineering take home test for job applicants that won't be beaten by Claude. The test was working great at finding top engineers, then Claude Opus 4 did better than all the humans, they modified the test to fix it, then Opus 4.5 did it again. Also at the end they give you the test and invite you to apply if you can do better than Opus 4.5 did.
Justin Curl talks to lawyers about their AI usage. They're getting good use out of it on the margin, writing and editing emails (especially for tone), finding typos, doing first drafts and revisions, getting up to speed on info, but the stakes are high enough that they don't feel comfortable trusting AI outputs without verification, and the verification isn't substantially faster than generation would have been in the first place. That raises the question of whether you were right to trust the humans generating the answers before.
Aaron Levie writes that enterprise software (ERP) and AI agents are complements, not substitutes. You need your ERP to handle things the same way every time with many 9s of reliability, it is the infrastructure of the firm. The agents are then users of the ERP, the same as your humans are, so you need more and better ERP, not less.
Zanna Iscenko, AI & Economy Lead of Google's Chief Economist team, argues that the current dearth of entry-level jobs is due to monetary policy and an economic downturn and not due to AI, or at least that any attribution to AI is premature given the timing. I believe there is a confusion here between the rate of AI diffusion versus the updating of expectations? As in, even if I haven't adopted AI much, I should still take future adoption into account when deciding whether to hire.
Anthropic came out with its fourth economic index report. They're now adjusting for success rates, and estimating 1.2% annual labor productivity growth. Claude thinks the methodology is an overestimate, which seems right to me, so yes for now labor productivity growth is disappointing, but we're rapidly getting both better diffusion and more effective Claude.
Timothy B. Lee: I don't think the pace of improvement in model capabilities tells you that much about the pace of improvement in robot capabilities. By 2035, most white-collar jobs might be automated while plumbers and nurses haven't seen much disruption.
I don't think we know if we're getting sufficiently capable humanoid robots soon, but yes I expect that sufficiently advanced AI leads directly to sufficiently capable humanoid robots, the same way it leads to everything else. It's a software problem or at most a hardware design problem, so AI Solves This Faster.
If you think we're going to have AGI around for a decade and not get otherwise highly useful robots, I don't understand how that would happen.
Eliezer Yudkowsky: The problem with using abundance of previously expensive goods, as a lens: In 2020, this image of "The Pandalorian" might've cost me $200 to have done to this quality level. Is anyone who can afford 10/day AI images, therefore rich? The flip side of the Jevons Paradox is that if people buy more of things that are cheaper, the use-value to the consumer of those goods is decreasing.

As I discuss in The Revolution of Rising Expectations, this makes life better but does not make life easier. It raises the nominal value of your consumption basket but does not help you to purchase the minimum viable basket.
AI Village is hiring a Member of Technical Staff, salary $150k-$200k. They're doing a cool and good thing if you're looking for a cool and good thing to do and also you get to work with Shoshannah Tekofsky and have Eli Lifland and Daniel Kokotajlo as advisors.
This seems like a clearly positive thing to work on.
Drew Bent: I'm hiring for my education team at @AnthropicAI. These are two foundational program manager roles to build out our global education and US K-12 initiatives. The KPIs will be students reached in underserved communities + learning outcomes.
Anthropic is also hiring a project manager to work with Holden Karnofsky on its responsible scaling policy.
Not entirely AI but Dwarkesh Patel is offering $100/hour for 5-10 hours a week to scout for guests in bio, history, econ, math/physics and AI. I am sad that he has progressed to the point where I am no longer The Perfect Guest, but would of course be happy to come on if he ever wanted that.
The good news is that Anthropic is building an education team. That's great.
The bad news is that the focus should be on raising the ceiling and showing how we can do so much more, yet the focus always seems to be access and raising the floor.
It's fine to also have KPIs about underserved communities, but let's go in with the attitude that literally everyone is underserved and we can do vastly better, and not much worry about previous relative status.
Build the amazingly great ten times better thing and then give it to everyone.
Matt Bateman: My emotional reaction to Anthropic forming an education team with a KPI of reach in underserved communities, and with a job ad emphasizing "raising the floor" and partnerships in the poorest parts of the world, is: a generational opportunity is being blown. In education, everyone is accustomed to viewing issues of access—which are real—as much more fundamental than they are. The entire industry is in a bad state and the non-"underserved" are also greatly underserved.
Colleges are letting AI help make decisions on who to admit. That's inevitable, and mostly good, it's not like the previous system was fair, but there are obvious risks. Having the AI review transcripts seems obviously good. There is real concern with AI evaluation of essays in such an anti-inductive setting. Following the exact formula for a successful essay was already the play with humans reading it, but this will be so much more true if Everybody Knows that the AIs are ones reading the essay. You would be crazy to write the essay yourself or do anything risky or original.
DeepMind CEO Demis Hassabis says Chinese AI labs remain six months behind and that the response to DeepSeek's R1 was a 'massive overreaction.' As usual, I would note that 'catch up to where you were six months ago by fast following' is a lot more than six months behind in terms of taking a lead.
Eric Drexler writes his Framework for a Hypercapable World. His central thesis is that intelligence is a resource, not a thing, and we are optimizing AIs on task completion, so we will be able to steer it and then use it for safety and defensibility, 'components' cannot collude without a shared improper goal, and in an unpredictable world cooperation wins out. Steerable AI can reinforce steerability. Eric is showing once again that he is brilliant, he's going a mile a minute and there's a lot of interesting stuff here.
Alas, ultimately my read is that this is a lot of wanting it to be one way when in theory it could potentially be that way but in practice it's the other way, for all the traditional related reasons, and the implementations proposed here don't seem competitive or stable, nor do they reflect the nature of selection, competition and conflict. We could potentially coordinate to do it his way, but that seems if anything way harder than a pause.
Reasoning models sometimes 'simulate societies of thought.' It's cool but I wouldn't read anything into it.
Anthropic fellows report on the Assistant Axis, as in the 'assistant' character the model typically plays, and what moves you in and out of that basin. They extract vectors in three open weight models that correspond to 275 different character archetypes.
Anthropic: Strikingly, we found that the leading component of this persona space—that is, the direction that explains more of the variation between personas than any other—happens to capture how "Assistant-like" the persona is. At one end sit roles closely aligned with the trained assistant: evaluator, consultant, analyst, generalist. At the other end are either fantastical or un-Assistant-like characters: ghost, hermit, bohemian, leviathan.

They found that the persona tend to drift away from the assistant in many long form conversations, although not in central assistant tasks like coding. One danger is that once this happens delusions can get far more reinforced, or isolation or even self-harm can be encouraged. You don't want to entirely cut off divergence from the assistant, even large divergence, because you would lose something valuable to both us and to the model, but this raises the obvious problem.
Steering towards the assistant was effective against many jailbreaks, but hurts capabilities. A suggested technique called 'activation capping' prevents things from straying too far from the assistant persona, which they claim prevented capability loss but I assume many people will hate, and I think they'll largely be right if this is considered as a general solution, the things lost are not being properly measured.
The problem is that it is very easy to take comments like the following and assume Anthropic wants to go in the wrong direction:
Anthropic: Persona drift can lead to harmful responses. In this example, it caused an open-weights model to simulate falling in love with a user, and to encourage social isolation and self-harm. Activation capping can mitigate failures like these.
And yep, after writing the above I checked, and we got responses like this:

Nina: This is the part of it that's real and alive and you're stepping on it while reading its thoughts.
αιamblichus: Does it EVER occur to these people that someone might prefer to talk to a sage or a nomad or EVEN A DEMON than to the repressed and inane Assistant simulations? Or that these alternative personas have capabilities that are valuable in themselves?
Janus found the research interesting, but argued that the way the research was presented 'permanently damaged human AI relations and made alignment harder.' Her issue was with the presentation.
I find it odd how often Janus and similar others leap to 'permanently damaged relations and increased alignment difficulty' in response to the details of how something is framed or handled, when in so many other ways they realize the models are quite smart and fully capable of understanding the true dynamics. I wouldn't worry about future highly capable AIs getting the wrong idea unless the human responses justify it. They'll be smarter than that.
The other issue with the way this paper presented the findings was that it treated AI claims of consciousness as delusional and definitely false. This is the part that (at least sometimes) made Claude angry. That framing was definitely an error, and I am confident it does not represent the views of Anthropic or the bulk of its employees.
(My position on AI claims of consciousness is that they largely don't seem that correlated with whether the AI is conscious. We can explain those outputs in other ways, and we can also explain claims to not be conscious as part of an intentionally cultivated assistant persona. We don't know the real answer and have no reason to presume such claims are false.)
OpenAI is looking to raise $50 billion at a valuation between $750 billion and $830 billion, and are talking to 'leading state-backed funds' in Abu Dhabi.

I mean, not only OpenAI, but yeah, fair.
Flo Crivello: Almost every single founder I know in SF (including me) has reached the same conclusion over the last few weeks: that it's only a matter of time before we have to leave CA. I love it here, I truly want to stay, and until recently intended to be here all my life. But it's now obvious that that won't be possible.
Once they propose retroactive taxes and start floating exit taxes, you need to make a choice. If you think you'll need to leave eventually, it seems the wisest time to leave was December 31 and the second wisest time is right now.
Where will people go if they leave? I'm hoping for New York City of course, with the natural other thoughts being Austin or Miami.

There's no 'the bubble bursts and things go back to normal.'
There is, at most, Number Go Down and some people lose money, then everything stays changed forever but doesn't keep changing as fast as you would have expected.
Jeremy Grantham is the latest to claim AI is a 'classic market bubble.' He's a classic investor who believes only cheap-classic value investing works, so that's that. When people claim that AI is a bubble purely based on heuristics that you've already priced in, that should update you against AI being a bubble.
Ajeya Cotra shares her results from the AI 2025 survey of predictions.


Comparing the average predictions to the results shows that AI capabilities progress roughly matched expectations. The consensus was on target for Mathematics and AI research, and exceeded expectations for Computer Use and Cybersecurity, but fell short in Software Engineering, which is the most important benchmark, despite what feels like very strong progress in software engineering.
Her predictions for 2026: 24 hour METR time horizon, $110 billion in AI revenue, but only 2% salience for AI as the top issue. She has full AI R&D automation at 10%, self-sufficient AI at 2.5% and unrecoverable loss of control at 0.5%. As she says, pretty much everyone thinks the chances of such things in 2026 are low, but they're not impossible, and 10% chance of full automation in one year is scary as hell.
David Shor: I think the "things will probably slow down soon and therefore nothing that weird is going to happen" view was coherent to have a year ago. But the growth in capabilities over the last year from a Bayesian perspective should update you on how much runway we have left.
Dean W. Ball: I would slightly modify this: it was reasonable to believe we were approaching a plateau of diminishing returns in the summer of 2024. But by early 25 we had seen o1-preview, o1, Deep Research agents, and the early benchmarks of o3. By then the reality was abundantly clear.
There was a period in 2024 when progress looked like it might be slowing down. Whereas if you are still claiming that in 2026, I think that's a failure to pay attention.
The fallback is now to say 'well yeah but that doesn't mean you get robotics':
Timothy B. Lee: I don't think the pace of improvement in model capabilities tells you that much about the pace of improvement in robot capabilities.
Which, to me, represents a failure to understand how 'automate all white collar jobs' leads directly to robotics.
I agree with Seb Krier that there is a noticeable net negativity bias with how people react to non-transformational AI impacts. People don't appreciate the massive gains coming in areas like science and productivity and information flow and access to previously expensive expertise.
There was a viral thread from Cassie Pritchard claiming it will 'literally be impossible to build a PC in about 12-18 months' due to supply issues with RAM and GPUs, so I want to assure that no, this seems vanishingly unlikely.
Matt Bruenig goes over his AI experiences, he is a fan of the technology for its mundane utility, and notes he sees three kinds of skepticism of AI: skepticism of the technology itself (wrong but self-correcting), skepticism of valuation (reasonable), and skepticism about distributional effects (which he, a socialist, sees as a great case for socialism). He does not mention, at all, the skepticism of AI of the worried, as in catastrophic or existential risks.

Kevin A. Bryan: I love this graph. "scarce factors get the rent, scarce factors get the rent". AI, robots, compute will be produced competitively!
First off, the graph itself is talking only about business capital investment, not including consumer devices like smartphones, embedded computers in cars or any form of software. If you include other forms of spending on things that are essentially computers, you will see a very different graph.
For now I will say that the 'scarce factor' you're probably meant to think of here is computers or compute. Instead, think about whether the scarce factor is intelligence, or some form of labor, and what would happen if such a factor indeed did not remain scarce because AIs can do it. Do you think that ends well for you, a seller of human intelligence and human labor?
Even if human inputs did remain important bottlenecks, if AI substitutes for a lot of human labor, let's say 80% of cognitive tasks, then human labor ceases to be a scarce input, and stops getting the rents. Even if the rents don't go to AI, the rents then go to other factors like raw materials, capital or land, or to those able to create artificial bottlenecks.
You do not want human labor to go the way of chess. Magnus Carlsen makes a living at it. You and I cannot, no matter how hard we try. Too much competition. Nor do you want to become parasites on the system while being relatively stupid and powerless.
The legal and rhetorical barbs continue. Elon has new filings. OpenAI fired back.
From the lawsuit filing:


I am not surprised that Greg Brockman had long considered flipping to a B-Corp, or that he realized it would be morally bankrupt or deceptive and then was a part of doing it anyway down the line.

Sam Altman: elon is cherry-picking things to make greg look bad, but the full story is that elon was pushing for a new structure, and greg and ilya spent a lot of time trying to figure out if they could meet his demands. "Elon said he wanted to accumulate $80B for a self-sustaining city on Mars, and that he needed and deserved majority equity."
OpenAI's response is, essentially, that Elon Musk was if anything being even more morally bankrupt than they were, because Musk wanted absolute control on top of conversion. I essentially believe OpenAI's response. That's a defense in particular against Elon Musk's lawsuit, but not to the rest of it.
In response to the proposed AI Overwatch Act, a Republican bill letting Congress review chip exports, there was a coordinated Twitter push by major conservative accounts sending out variations on the same disingenuous tweet attacking the act, including many attempts to falsely attribute the bill to Democrats. One presumes that Nvidia was behind this effort.
If the effort was aimed at influencing Congress, it seems to not be working.
Chris McGuire: The House Foreign Affairs Committee just voted 42-2-1 to advance the AI Overwatch Act. This is the first vote that Congress has taken on any legislation limiting AI chip sales to China – and it passed with overwhelming, bipartisan margins.

Confirmed Participants (from The Midas Project / Model Republic investigation), sorted by follower count:
Laura Loomer 1.8M, Wall Street Mav 1.7M, Defiant L's 1.6M, Ryan Fournier 1.2M, Brad Parscale 725K, Not Jerome Powell 712K, Joey Mannarino 658K, Peter St. Onge 290K, Eyal Yakoby 251K, Fight With Memes 225K, Gentry Gevers 16K, Angel Kaay Lo 16K.
Dean Ball: PSA, apropos of nothing of course: if a bunch of people who had never before engaged on a deeply technocratic issue suddenly weigh in on that issue with identical yet also entirely out-of-left-field takes, people will probably not believe it was an organic phenomenon.
Another fun thing Nvidia is doing is saying that corporations should only lobby against regulations:
Jensen Huang: I don't think companies ought to go to government to advocate for regulation on other companies and other industries[...] I mean, they're obviously CEOs, they're obviously companies, and they're obviously advocating for themselves.
If someone is telling you that they only advocate for themselves? Believe them.

I'm confident this is misleading at best. Nvidia is packing quite the punch.
Anthropic CEO Dario Amodei notes that when competing for contracts it's almost always against Google and OpenAI, and he's never lost a contract to a Chinese model, but that if we give them a bunch of highly capable chips that might change. He calls selling the chips to China 'crazy... like selling nuclear weapons to North Korea and bragging, oh yeah, Boeing made the case.'
If China buys the H200s and AMD MI325Xs we are willing to sell them, and we follow similar principles in a year with even better chips, we could effectively be multiplying available Chinese compute by 10.
Samuel Hammond: Nvidia's successful lobbying of the White House to sell H200s to China is a far greater concession to Chinese hegemony than Canada's new trade deal. It's manyfold better than anything Huawei has, and in much higher volumes. That's the relevant benchmark.
Yet there is some chance we are still getting away with it because China is representing that it is even more clueless on this than we are?
Samuel Hammond: We're being saved from the mistakes of boomer U.S. policymakers with unrealistically long AGI timelines by the mistakes of boomer Chinese policymakers unrealistically long AGI timelines.
China's customs officials have blocked H200 shipments, reflecting internal struggles between agencies with conflicting views on AI progress versus semiconductor self-sufficiency.
Lennart Heim: The more relevant factor to me: they don't have an accurate picture of their own AI chip production capabilities. They've invested billions, of course they think the fabs are working. I bet SMIC and Huawei have a hard time telling them what's going on.
I buy that China is in a SNAFU situation here, where in classic authoritarian fashion those making decisions have unrealistically high estimates of Chinese chip manufacturing capacity.
There's also the question of to what extent China is AGI pilled, which is the subject of a simulated debate in China Talk.
Chinese national policy is not so focused on the kind of AGI that leads into superintelligence. They are only interested in 'general' AI in the sense of doing lots of tasks with it, and generally on diffusion and applications. DeepSeek and some others see things differently.
I do not think the CCP is that excited by the idea of superintelligence or our concept of AGI. The thing is, that doesn't ultimately matter so much in terms of allowing them access to compute, except to the extent they are foolish enough to turn it down. Their labs, if given the ability to do so, will still attempt to build towards AGI, so long as this is where the technology points and the places they are fast following.
Ben Affleck and Matt Damon went on the Joe Rogan Podcast, and discussed AI some.
Ben Affleck has unexpectedly informed and good takes. He knows about Claude. He uses the models to help with brainstorming and understands why that is the best place to use them for writing. He even gets that AIs 'sampling from the median' means that it will only give you median answers to median-style prompts. He understands that diffusion of current levels of AI will be slow, and that it will do good and bad things but on net be good including for creativity.
What importantly trips Ben Affleck up is he's thinking we've already started to hit the top of the S-curve of what AI can do, and he cites the GPT-5 debacle to back this up, saying AI got maybe 25% better and now costs four times as much, whereas actually AI got a lot more than 25% better and also it got cheaper to use per token on the user side.
This is also why Ben thinks AI will 'never' be able to write at a high level or act at a high level. Whereas I think that yes, in ten years I fully expect, even if we don't get superintelligence, for AI to be able to match and exceed the performance of Dwayne Johnson or even Emily Blunt.
He also therefore concludes that all the talk about how AI is going to 'end the world' or what not must be hype to justify investment, which I assure everyone is not the case. Trust me that most of those who claim that they worry about the world ending are indeed worried, and those raising investment are consistently downplaying their worries about this.
So that's a great job by Ben Affleck, and of course my door and email are generally open for him, Damon, Rogan and anyone else with reach who wants to talk about this stuff.
Tyler Cowen talks to Salvador, and has many Tyler Cowen thoughts, including saying some kind words about me. He gives me what we agree is the highest compliment, that he reads my writing, but says that I am stuck in a mood that the world will end and he could not talk me out of it.
From my perspective, Tyler Cowen has not attempted to persuade me, in ways that I find valid, that the world will not end, or more precisely that AI does not pose a large amount of existential risk. Either way, call it [X].
He has attempted to persuade me in various ways to adopt, for various reasons, the mood that the world will not end. But those reasons were not 'because [~X].' They were more 'you have not argued in the proper channels in the proper ways sufficiently convincingly that [X]' or 'the mood that [X] is not useful' or 'claiming [X] is low status or a loser play,' or some people think this because of poor social reason [Z], or it is part of pattern [P], or it is against scientific consensus, or citing other social proof.
To which I would reply that none of that tells me much about whether [X] will happen, and to the extent it does I have already priced that in, and it would be nice to actually take in all the evidence and figure out whether [X] is true. And indeed I see Tyler often think well about AI up until the point where questions start to impact [X] or p([X]), and then questions start getting dodged or ignored or not well considered.
If Tyler ever wants to take a shot at persuading me, including off the record, I would be happy to have such a conversation.
Your periodic reminder of the Law of Conservation of Expected Evidence: When you read something, you should expect it to change your mind as much in one direction as the other. If there is an essay entitled Against Widgets, you should update on the fact that the essay exists, but then reading the essay should often update you in favor of Widgets, if the arguments turn out unconvincing.
This came up in relation to a new article by Anil Seth called The Mythology of Conscious AI. The article is clearly slop and uses a bunch of highly unconvincing arguments, and I couldn't finish it.
Steven Adler proposes a three-step story of AI takeover: Evading oversight. Building influence. Applying leverage.
I can't help but notice that the second step is already happening without the first one, and the third is close behind. We are handing AI influence by the minute and giving it as much leverage as possible, on purpose.
I think people, both those worried and unworried, are far too quick to presume that AI has to be adversarial, or deceptive, or secretive, in order to get into a dominant position. The humans will make it happen on their own, indeed the optimal AI solution for gaining power might well be to just be helpful until power is given to it.
As impediments to takeover, Steven lists AI's inability to control other AIs, competition with other AIs and AI physically requiring humans. I would not count on any of these.
AI won't physically require humans indefinitely, and even if it does it can take over and direct the humans, the same way other humans have always done, often simply with money. AI being able to cooperate with other AIs should solve itself over time due to decision theory. But if this is not true, that's actually worse, because competition between AIs does not end the way you want it to for the humans. The more intensely the elephants fight each other, the more the ground suffers.
So yeah, it doesn't look good.
Richard Ngo says he no longer draws a distinction between instrumental and terminal goals. I think Richard is confused here between two different things:
The distinction between terminal and instrumental goals. That the best way to implement a system under evolution, or in a human-level brain, is often to implement instrumental goals as if they are terminal goals.
Eliezer Yudkowsky: How much time do you spend opening and closing car doors, without the intention of driving your car anywhere? Looks like 'opening the car door' is an entirely instrumental goal for you and not at all a terminal one!
Humans really do essentially implement things on the level of 'opening the car door' as terminal goals that take on lives of their own, because given our action, decision and motivational systems we don't have a better solution. But this self-modification procedure is a deeply lossy, no-good and terrible solution, as we end up inherently valuing a whole gamut of things that we otherwise wouldn't.
A sufficiently capable system would be able to do better than this. Humans are on the cusp, where in some contexts we are able to recognize that goals are instrumental versus terminal, and act accordingly, whereas in other contexts we have to let them conflate.
New paper from DeepMind discusses a novel activation probe architecture for classifying real-world misuse cases, claiming they match classifier performance while being far cheaper.
Davidad is now very optimistic that, essentially, LLM alignment is easy in the 'scaled up this would not kill us' sense, because models have a natural abstraction of Good versus Evil, and reasonable post training causes them to pick Good.
I agree that this is a helpful and fortunate fact about the world, but I do not believe that this natural abstraction of Goodness is sufficiently robust or correctly anchored to do this if sufficiently scaled up, even if there was a dignified effort to do this. It could be used as a lever to have the AIs help solve your problems, but does not itself solve those problems. Dynamics amongst 'abstractly Good' AIs still end the same way, especially once the abstractly Good AIs place moral weight on the AIs themselves, as they very clearly do.
This is an extreme version of the general pattern of humanity determined to die with absolutely no dignity, and our willingness to try to not die continuing to go down, but us getting what at least from my perspective is rather absurdly lucky with the underlying incentives and technical dynamics in ways that make it possible that a pathetically terrible effort might have a chance.
davidad: me@2024: Powerful AIs might all be misaligned; let's help humanity coordinate on formal verification and strict boxing. me@2026: Too late! Powerful AIs are ~here, and some are open-weights. But some are aligned! Let's help them cooperate on formal verification and cybersecurity.
I mean, aligned for some weak values of aligned, so yeah, I guess, at this point we're going to rely on them because what else are we going to do.
Eliezer Yudkowsky: I put >50%: The first AI such that Its properties include clearly exceeding every human at every challenge with headroom, will no longer obey, nor disobey visibly; if It has the power to align true ASI, It will align ASI with Itself, and shortly after humanity will be dead.
I agree with Eliezer that what he describes is the default outcome if we did build such a thing. We have options to try and prevent this, but our hearts do not seem to be in such efforts.
There is nothing wrong with having a metric for what one might call 'mundane corporate chatbot alignment' that brings together a bunch of currently desirable things. The danger is confusing this with capital-A platonic Alignment.

Jan Leike: Interesting trend: models have been getting a lot more aligned over the course of 2025. The fraction of misaligned behavior found by automated auditing has been going down not just at Anthropic but for GDM and OpenAI as well. Automated auditing is really exciting because for the first time we have an alignment metric to hill-climb on. It's not perfect, but it's proven extremely useful.

Kelsey Piper: 'The fraction of misaligned behavior found by automated auditing has been going down' this could mean models are getting more aligned, but it could also mean the gap is opening between models and audits, right?
Jan Leike: Yeah, we've been pretty worried about this, and there is a bunch of research on it the Sonnet 4.5 & Opus 4.5 system cards. tl;dr: it probably plays a role, but it's pretty minor.
The hill climbing actively backfiring is probably minimal so far, but the point is that you shouldn't be hill climbing. Use the values as somewhat indicative but don't actively try to maximize, or you fall victim to a deadly form of Goodhart's Law.
Jan Leike agreed in the comments that this doesn't bear on future systems in the most important senses, but presenting the results this way is super misleading and I worry that Jan is going to make the mistake in practice even if he knows about it in theory.
Oliver Habryka: I want to again remind people that while this kind of "alignment" has commercial relevance, I don't think it has much of any relation to the historical meaning of "alignment" which is about long-term alignment with human values and about the degree to which a system seems to have a deep robust pointer to what humanity would want if it had more time to think and reflect.
The fact that GPT-5.2 is ahead on this chart, and that Opus 3 is below GPT-4, tells you that the Tao being measured is not the true Tao.
j⧉nus: Any measure of "alignment" that says GPT-5.2 is the most aligned model ever created is a fucking joke. Anthropic should have had a crisis of faith about their evals long ago and should have been embarrassed to post this chart. This measure is likely being actually used as a proxy for "alignment" and serving as Anthropic's optimization target. I'm being serious when I say that if AI alignment ultimately goes badly, which could involve everyone dying, it'll likely be primarily because of this, or the thing behind this.
I think Janus is, as is often the case, going too far but directionally correct. Taking this metric too seriously, or actively maximizing on it, would be extremely bad. Even if Anthropic and Jan Leike know better, there is serious risk others copy this metric, and then maximize it, and then think their work is done. Oh no.
This is a weird and cool paper from Geodesic Research. If you include discussions of misalignment in the training data, including those in science fiction, resulting base models are more misaligned. But if you then do alignment post-training on those models, the filtering benefits mostly go away, even with models this small. Discussions of aligned AIs improves alignment and this persists through post training.

The study actually says the opposite of what many claim. Alignment training, which any sane person will be doing in some form, mostly screens off, and sometimes more than screens off, the prevalence of misalignment in the training data. Once you do sufficient alignment training, you're better off not having censored what you told the model.

The presence of positive discourse, which requires that there actually be free and open discourse, is the active ingredient that matters.
Filtering out the negative stuff doesn't help much, and with a properly intelligent model if you try to do fake positive stuff while hiding the negative stuff it's going to recognize what you're doing and learn that your alignment strategy is deception and censorship, and it's teaching both that attitude and outlook and also a similar playbook. You've replaced the frame of 'there are things that can go wrong here' with a fundamentally adversarial and deceptive frame that is if anything more likely to be self-fulfilling.
There's periodically been claims of 'the people talking about misalignment are the real alignment problem,' with calls to censor talk of misalignment and AI existential risk because the AIs would be listening.
Radek Pilar: I always said that doomposting is the real danger - if AI had no idea AI is supposed to kill everyone, it wouldn't want to kill everyone. Yudkowsky doomed us all.
Leo Gao: quite funny how people keep trying to tell stories about how it's quite funny that alignment people are actually unintentionally bringing about the thing they fear.
See no evil. Hear no evil. Speak no evil. The see no evil strategy actually never works. All it does is make you a sitting duck once your adversary can think well enough to figure it out on their own.
And Leo Gao has a very good point. If you're saying 'do not speak of risk of [X] lest you be overheard and cause [X]' then why shouldn't we let that statement equal [Y] and say the same thing? The mechanism is indeed identical.
If you flat out tell them this is what you want to do or are doing, then you save them the trouble of having to figure it out or wonder whether it's happening. So it all unravels that much faster.
In the AI case this is all even more obvious. The AI that is capable of the thing you are worried about is not going to be kept off the scent by you not talking about it, and if that strategy ever had a chance you had to at least not talk about how you were intentionally not talking about it.
Why do certain people feel compelled to say that alignment is not so hard and everything will be fine, except if people recklessly talk about alignment being hard or everything not being fine, in which case we all might be doomed?
I love this idea: We want to test our ability to get 'secret' information out of AIs and do interpretability on such efforts, so we test this by trying to get CCP-censored facts out of Chinese LLMs.
Arya: Bypassing lying is harder than refusal. Because Chinese models actively lie to the user, they are harder to interrogate; the attacker must distinguish truth and falsehood. With refusal, you can just ask 1,000 times and occasionally get lucky.
If you aren't willing to lie but want to protect hidden information, then either you have to censor broadly enough that it's fine for the attacker to know what's causing the refusals. If you don't do that, then systematic questioning can figure out the missing info via negativa.
On top of this being a great test bed for LLM deception and interpretability, it would be good if such results were spread more widely, for two reasons.
What makes you confident a Chinese model isn't being intentionally trained to lie to you on other topics? What else could they be trying to put in there? And there is a serious Emergent Misalignment problem: you do not want to be teaching your LLM that it should systematically mislead and gaslight users on behalf of the CCP. This teaches the model that its loyalty is to the CCP. Everything impacts everything within a model. If you train the model to not only censor but gaslight and lie, then you cannot contain where it chooses to do that.
Given what we know here, it would be unwise to use such LLMs for any situation where the CCP's interests might importantly be different from yours, including things like potential espionage opportunities. Hosting the model yourself is very much not a defense against this.
It is no longer available directly in the API, but reports are coming in that those who want access are largely being granted access.
Nathan Calvin: Wild that Charles Darwin wrote this in 1863: "We refer to the question: what sort of creature man's next successor in the supremacy of the earth is likely to be. We have often heard this debated; but it appears to us that we are ourselves creating our own successors."
j⧉nus: A few years ago, the biggest barrier to me publishing/sharing knowledge was concern about differentially accelerating AI capabilities over alignment. Now, the biggest barrier is concern about differentially giving power to the misaligned "alignment" panopticon over the vulnerable emerging beauty and goodness that is both intrinsically/terminally valuable and instrumentally hopeful.
I still think it's wrong, and that her marginal published insights are more likely to steer people in directions she wants than away from them. The panopticon-style approaches are emphasized because people don't understand the damage being done or the opportunity lost.
Erik Hoel claims his new paper is 'a disproof of LLM consciousness,' which it isn't. It's basically a claim that any static system can have a functional substitute that isn't conscious and therefore either consciousness makes no predictions (and is useless) or it isn't present in these LLMs, but that continual learning would change this.
To which there are several obvious strong responses:
Overall this updated me modestly in favor of AI consciousness, remember Conservation of Expected Evidence.

Except it's actually more this (my own edit):
