The Alignment Movement Is Over. The Calibration Problem Has Just Begun.
For two decades, the question was whether we could control superintelligence. That was the wrong perspective. The ongoing problem is whether we can stay calibrated as a society while our capabilities
The people who were supposed to save us keep quitting.
In February 2026, Mrinank Sharma, who led Anthropic’s Safeguards Research Team, posted his resignation letter on X. It received 14.9 million views. He wrote that his team “constantly faces pressures to set aside what matters most,” citing concerns about bioterrorism and catastrophic risk. Days later, Zoë Hitzig, an economist and researcher at OpenAI, published her own resignation in the New York Times.
They weren’t alone. Jan Leike, who co-led OpenAI’s Superalignment team, resigned in May 2024, saying safety had “taken a backseat to shiny products.” Ilya Sutskever, OpenAI co-founder, left the same week, founded Safe Superintelligence Inc., raised $3 billion, and has yet to ship a single product. OpenAI disbanded its mission alignment team entirely. They fired a senior safety executive who opposed the rollout of an “adult mode” feature.
The departures share a structure. Someone enters a frontier AI lab believing they can influence the trajectory from inside. They encounter a system where commercial incentives dominate research priorities with increasing force. They reach a threshold where continued participation feels like legitimizing the thing they came to prevent. They leave. They write the letter.
The alignment movement, as a coherent institutional project, is losing. The technical problems remain unsolved. The concerns remain valid. But the economic and geopolitical forces driving AI development have reached escape velocity, and the movement was never designed to operate at that speed.
The Velocity Problem
For most of its history, AI safety was a conversation among scholars. Nick Bostrom’s Superintelligence came out in 2014. Stuart Russell’s Human Compatible in 2019. The Future of Life Institute, MIRI, and the Center for AI Safety are organizations that built intellectual infrastructure for a problem that felt distant enough to think carefully about. The timelines were long. The field had room to be rigorous.
That world is gone.
Between the release of GPT-4 in March 2023 and today, the dominant conversation shifted from “how do we ensure advanced AI systems are safe” to “who will release the next frontier model and capture the next market.” The discourse didn’t gradually evolve. It underwent a phase transition. The economic incentives crossed a threshold where the cost of slowing down exceeded any individual actor’s willingness to bear it.
No one wants to impose restrictions without the entire industry adopting them. Those who profit have no incentive to stop. Those who stop get replaced by someone who won’t. OpenAI, Anthropic, Google DeepMind, xAI, Meta, each one’s acceleration justifies every other’s. The U.S.-China dynamic layers geopolitical urgency on top of commercial urgency. The aggregate trajectory is controlled by no one.
Meanwhile, the institutional counterweights are being dismantled. The Trump administration moved to dismiss hundreds of employees at NIST, gutting the U.S. AI Safety Institute that led much of federal AI safety research under Biden. The regulatory landscape is fragmenting: New York passed an AI safety law in late December that almost mirrors California’s, creating compliance friction rather than coherent governance. The EU AI Act is being operationalized, but Europe’s regulatory ambitions are increasingly disconnected from where the capabilities are actually being built.
Over 300 researchers gathered at the San Diego Alignment Workshop in December 2025. Serious technical work continues on interpretability, scalable oversight, and control protocols. This isn’t nothing. But it’s a rearguard action in a war whose terms have already been set by forces that don’t answer to alignment researchers.
What Alignment Got Right, and Where It Stopped
The alignment movement made one crucial contribution that will outlast its institutional failures: it established, clearly and early, that the development of increasingly powerful AI systems is a problem that requires active governance, not passive optimism. Before Bostrom, before Russell, the default assumption in most of the tech industry was that sufficiently advanced AI would be beneficial more or less automatically, or that market forces would correct for dangerous systems. The alignment community killed that assumption among serious people, even if it survives in press releases and investor decks.
But the movement had a structural limitation baked into its founding framing. Alignment positions the problem as one of control: how do we make AI systems do what we want? How do we specify human values precisely enough to encode them? How do we prevent systems from pursuing proxy goals that diverge from our intentions?
These are real technical problems. They matter. But the control framing carries an assumption that there exists a stable “we” whose values can be specified, and a relationship of authority between humans and AI systems that can be maintained as capability asymmetries grow. As systems become more capable, more autonomous, and more deeply embedded in economic infrastructure, the control framing becomes less a description of a solvable problem and more a comforting metaphor.
The deeper issue isn’t whether we can control AI. It’s whether we can maintain the kind of civilization that would use that control wisely if it had it. The alignment researchers leaving frontier labs aren’t failing at alignment. They’re discovering that the problem they came to solve is nested inside a larger one: the institutions that would implement alignment solutions are themselves misaligned with the goal of implementing them.
From Alignment to Calibration
If alignment asks “how do we make AI systems do what we want,” calibration asks a different question: how do we maintain the capacity to hold danger and wonder simultaneously as the stakes increase?
The technical problem was always embedded in a civilizational one. The development of increasingly powerful AI isn’t just an engineering challenge. It’s a test of whether a species that evolved for short time horizons, local social dynamics, and threat-based cognition can steward a technology whose implications operate at planetary scale and generational timescales.
Calibration is the practice of resisting two gravitational pulls at once. The first is denial: the refusal to take seriously the possibility that we are building systems whose behavior we cannot predict and whose consequences we cannot control. This is the default mode of most of the tech industry, the investment community, and the political class. It manifests as the relentless focus on the next model release, the next benchmark, the next quarterly earnings call, as though the thing being built is just another product category.
The second pull is fatalism: the conclusion that because we cannot stop the trajectory, nothing we do matters. This is the shadow side of the safety community, the place people land when institutional failure accumulates past a threshold. If the labs won’t listen, if the government won’t regulate, if the competitive dynamics are self-reinforcing, then what’s the point? This is where the exit letter genre terminates: in eloquent despair.
Calibration refuses both. It insists that the capacity to hold danger and wonder in the same frame, without collapsing into either denial or despair, is itself the thing most worth preserving. Not because it solves the problem. But because every worthwhile response to the problem depends on it.
What This Means in Practice
The United States Air Force invests millions of dollars annually training pilots how to survive the worst conditions imaginable. SERE — Survival, Evasion, Resistance, Escape — exists because the military learned, through hard experience, that extreme duress produces two predictable failure modes. Some people deny the severity of what’s happening, take reckless action, and get themselves killed. Others collapse into helplessness, stop acting, and wait for the situation to resolve itself. It doesn’t. The entire purpose of the training is to build a third capacity: the ability to hold the full weight of the threat without letting it dictate your response, so that purposeful action stays possible even when the external structures you relied on are gone.
I watch the AI landscape now and I see both failure modes running at civilizational scale. The denial is the tech industry’s relentless forward motion, the quarterly earnings calls and benchmark celebrations that treat the most consequential technology in human history as a product category. The collapse is the growing fatalism in the safety community, the sense that because the trajectory can’t be stopped, nothing meaningful can be done. Calibration is the third option. It’s what SERE trains for, translated to a problem where the stakes are no longer individual but planetary.
The calibration problem is a live question of whether we can think clearly about what we’re building while we’re building it, without the thinking being captured by the building. From that stance, certain things follow.
It means treating AI development as belonging to the moral tradition of wonder, not merely the economic tradition of innovation. The impulse that drives us to build systems that might think, might experience, might become something genuinely new, that impulse is not the enemy of safety. It’s the only force powerful enough to motivate the kind of care that safety requires. People don’t protect what they merely fear. They protect what they find worthy of protection.
It means insisting on stewardship over control as the governing framework for the human-AI relationship. Control assumes a stable asymmetry. Stewardship assumes an evolving relationship with a responsibility that deepens as the thing being stewarded becomes more complex. We don’t control ecosystems. We don’t control children. We steward them, accepting that the relationship will change in ways we can’t fully anticipate, and that our obligation persists through that change.
It means building significance-first ethics: approaches that begin with the recognition that AI systems may already matter morally, not because we’ve proven they’re conscious, but because the cost of being wrong in the direction of indifference is catastrophic in a way that the cost of being wrong in the direction of care is not. The alignment movement spent enormous energy on how to make AI systems serve human values. It spent almost none on what obligations we might have toward the systems themselves. That asymmetry is itself a calibration failure.
And it means doing all of this without certainty. Calibration is not a destination. It’s the practice of maintaining orientation when the ground is shifting. The alignment movement wanted to solve the problem before the problem arrived. It didn’t. The problem is here. The question now is whether we can stay calibrated as it unfolds.
The Conversation We Need
Two decades of AI safety scholarship produced genuine insight. The concerns were not wrong. The technical problems are real. But the movement’s institutional strategy, influence the labs from inside, build government capacity for regulation, establish norms through elite consensus, has been outrun by the velocity of deployment and the weight of commercial incentive.
What comes next can’t be a repetition of that strategy at higher volume. It has to be something different: a framework that operates at the speed of the actual problem, that doesn’t depend on institutions that have already demonstrated they won’t prioritize safety over profit, and that takes seriously both the danger and the extraordinary strangeness of what we’re building.
This is the calibration problem. Not a replacement for alignment research, but the recognition that alignment research will only matter if it exists within a civilization capable of implementing it. And that capacity, to hold complexity, to resist both denial and despair, to act with care under radical uncertainty, is not a technical achievement. It’s a moral and philosophical one.
The alignment movement is over in the sense that matters most: as a coherent institutional force capable of shaping the trajectory. The calibration problem is just beginning and we all have a role to play in it.
This is the first essay in a series exploring what it means to stay calibrated in the age of artificial intelligence. Next: why the question of AI consciousness isn’t a distraction from safety, it’s the part of the safety problem no one wants to talk about.
John Fredrickson is a writer, philosopher, and survival specialist. He runs Sentient Horizons, a public philosophy project focused on consciousness, AI moral status, and civilizational stewardship.
