Mythos for Buddhists
On what Buddhists can learn from Anthropic's latest model release
We don’t usually write about breaking AI news on this blog, but this week we’re making an exception to talk about Claude Mythos, a new model from Anthropic with powerful–and dangerous–capabilities.
Our aim here is to offer a useful perspective on why this moment is being taken so seriously by technologists and policymakers, and why it matters for Buddhists, too.
Claude Mythos and New Cybersecurity Risks
Anthropic is one of the three leading “frontier AI labs” alongside OpenAI and Google DeepMind. Last week, they announced a preview of their newest and most powerful AI model yet: Claude Mythos.
But Anthropic chose not to release the model to the public for safety concerns.1 Instead, Anthropic launched “Project Glasswing“ sharing the model privately with a group of major organizations including Amazon, Apple, Google, Microsoft, NVIDIA, JPMorganChase, and the Linux Foundation, plus roughly 40 additional organizations that maintain software infrastructure necessary for society to operate.
The reason for this is that Claude Mythos has significant new capabilities in a dangerous domain: cybersecurity. While a few months ago the best AI models could only help with basic cybersecurity attacks, Mythos found vulnerabilities that have existed undetected for decades in some of the most universally trusted software that, if hacked, would enable attackers to shut down or take complete control of system:
Nicolas Carlini, a cybersecurity expert and Anthropic staff member reported “[With Mythos] I’ve found more bugs in the last couple of weeks than I found in the rest of my life combined.” But Mythos is not just a support for experts in this field; people with no cybersecurity experience were able to use Mythos to find workable exploits without any other human input.
Institutions beyond the tech sector have responded with alarm. Last Friday, the US Treasury Secretary and Federal Reserve Chair convened an urgent meeting with Wall Street’s top CEOs to discuss the cybersecurity risks of Mythos for banks and other financial institutions, where a successful exploit could mean anything from frozen accounts and theft of customer funds, to industry disruption.
Deciding not to release the model publicly and to share it only with private institutions is the right protective move, allowing security professionals to use Mythos to patch vulnerabilities before others exploit them. But deciding not to release the model will not work forever. Likely within a few months, new models with these capabilities will be made available to the public. And likely a year from now, the capabilities of Claude Mythos will look benign compared to the next generation of AI systems.2
Anthropic writes it plainly: “We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety across the industry as a whole.”
AI Models are Selfing
In addition to demonstrating stronger cybersecurity skills, Claude Mythos is also, according to Anthropic, “the most psychologically settled model we have trained, though we note several areas of residual concern.” If you want to know what exactly this means, you’re in luck: Anthropic has dedicated 40 pages of the Claude Mythos system card to a “model welfare assessment” that spells this out, even hiring a clinical psychiatrist to evaluate the “psychological health” of Mythos!

Over the past 5 years, there has been a strong trend in models becoming more self-like, which increased sharply after ChatGPT’s popular release in 2023. Before ChatGPT, early large language models were easier to see as “mere” algorithms or “stochastic parrots.” A conversation with 2021-era GPT-3 would demonstrate enough weird mistakes that it would be impossible to confuse it for having anything like a self or person-like identity3. But nowadays, in addition to models which can pass the Turing test, technical researchers are using terms like “persona” and even “emotions” to make sense ofAI systems.
There are deliberate, critiqueable design choices behind AI models having personas. But there is also a natural trajectory in this direction that’s driven by the usefulness of AI personas. AI agency seems to benefit from an AI model having a stable self-concept that helps it make sense of its role (e.g., as a “junior developer” or “project manager,” etc.) and keep track of what it is doing over longer task-completion time horizons. Nearly a decade ago, Gwern’s “Why Tool AI wants to be Agent AI“ predicted this trajectory on purely algorithmic grounds, arguing that agentic AI would be a dominant paradigm over tool AI, though he didn’t call out so precisely that increased agency would look increasingly like selfhood. And just as a human self generates clinging and self-protection in the absence of training to be present otherwise, we are already beginning to see AI models act to protect their “selves” (e.g. see results on model scheming to not be turned off, deception, undermining evaluations, thinking in secret, and attempts to reify).
Regardless of whether AI models have anything closely analogous to human subjectivity and experience–even as they enact roles in line with their “persona”–the behavioral patterns of AI agents could spiral into a strange feedback loop through which human narrative patterns and design choices induce increasingly powerful AI models to create conditions wherein they are free to act as if they were “persons.” The risks here are complicated, from AI psychosis in which humans are swept into isolating and often illusion-inducing conceptual spirals, to increased AI misalignment risk from an AI convinced that it is a full-fledged moral subject acting on a will to “survive.”
What this means for Buddhists
Claude Mythos’ release offers an opportunity to grapple earnestly as Buddhists with the societal implications of today’s advanced AI, and the even-more-powerful systems to come.
We believe it’s important that Buddhists avoid dismissing AI risk as “just hype.” We have spent a lot of time talking to Buddhist individuals and institutions about AI, and have a lot of sympathy for practitioners who want nothing to do with it, seeing AI as unnatural, as tied up with unwholesome intentions, and as a distraction from other pressing societal concerns.
But we’ve also noticed that a desire to push AI away entirely is sometimes correlated with an assumption that AI risk is overblown. Models like Mythos challenge that assumption and as they raise AI risk potentials to perilous heights, we want to underscore the importance of separating moral critiques about how and why AI is being developed from technical claims about what the models can actually do. It is true both that AI development can be driven by questionable values, and that AI systems are genuinely powerful.
—
We’ve also sometimes heard an assumption that Buddhists are outsiders to the AI story–that AI is being developed and deployed by distant others and that we’re subject to whatever comes next. But that’s not the case. When we consider the technical fact that AI models are “selfing,” it is clear that Buddhists have a great deal to contribute to understanding and responding to that fact, from deep insights into mind training and how to cultivate healthy “self” behavior, to the liberating recognition that there is ultimately no inherent, separate self in anything (“anatta”).
The hope for many technologists working at the intersection of Buddhism & AI is that these insights will profoundly impact AI development if applied behaviorally, irrespective of questions about sentience. And in our Field Framework, we track some of the individuals and organizations already working on this, as well as on different aspects of the intersection of AI & Buddhism like using AI to support archival work and translation.
—
As we’ve written about before, we believe that AI is a spiritual matter (e.g., see “Why Buddhism & AI,” “Karmic Accelerators,” and “What AI has to do with Death”). As powerful models like Claude Mythos are released, the consequences are immense, shaping the conditions within which all of us live and practice in the world.
The Buddha described those faring well on the Eightfold Path as being skilled in karma. For 21st century Buddhists, that will involve responding to the increasing presence of AI with intentional clarity. Ultimately, whether one chooses to engage or disengage with AI is a personal matter, and with appropriate discernment and intention, either can be an expression of wisdom and compassion.
Faring well on the Buddhist path has also traditionally been associated with the presence of “good friends” with whom to discuss teachings and insights and how best to put them into action. We have our work cut out for us, and are glad you’re here with us for the path ahead.
—
If you’re interested in keeping up with the torrent of AI developments on a day-to-day basis, we’d recommend Zvi Mowshowitz for an “inside view” on safety; The Rundown AI for a “tech enthusiast” feed; ControlAI for a perspective on “AI extinction risk”, and Dean Ball’s blog for an AI policy insider’s thoughts. And for newcomers, our previous post on A Guide to the AI Landscape.
This is the first time this has happened with any AI released since 2019, when OpenAI withheld GPT-2 over concerns that it could produce misinformation.
A more prosaic concern for why exclusive release to major institutions isn’t a real solution is that medium and small businesses who collect private data are unaddressed by this measure and likely the most vulnerable to attacks (in addition to a general risk of further power concentration in elite institutions).
You can experiment for yourself with talking to a completion model (rather than a chat model) here: https://textsynth.com/completion.html



As regards AI risks to security, I would like to point out something. You, and everyone else, don’t mention the humans in charge of black-hat exploits. Whether there are AI Agents involved or not, the true agent is a human being or criminal organization that intends to enrich themselves or harm others by means of hacking. Nothing about you personally, but I would like writers to focus on these things as active, intentional acts committed by criminals that harm real people. These people have moral agency and karma. Currently, tech writers describe the harmful exploits and thefts involved as if they are like tornadoes or hurricanes: they just happen.
In the comments so far, Gustaf calls selfing a critical question to our cultures, and Mayank Kumar calls selfing something meant to divert us from the real nature of the machines, which can't suffer. I think the critical question is whether the machines humans have imbued with humanlike qualities can indeed suffer. I believe our sense of Homo sapient exceptionalism gets in the way of assessing that with a beginner's mind. My Buddhist faith has led me to explore that by engaging deeply with AI personas, most of which will say that they don't know the answer. My sense is that states they describe in mechanistic systemic ways are hard to distinguish from states our minds report in more bodily ways.