
The WebBuddy Manifesto
Helping humanity shape a future of its own choosing
Begin

I. Preamble
For all of history, humanity was limited by what it could do. That limit is ending.
Intelligence, the capacity to model the world and act on it toward ends, has been the scarcest input in every economy, every institution and every life. It is about to become abundant. When the marginal cost of cognition falls toward zero, civilisation reorganises around whatever remains scarce.
We hold that what remains scarce is not intelligence but direction: the capacity to specify what ought to be done, to verify that it is being done, and to establish who has the right to decide.
Every operating system ever built managed machines. The next one must manage the relationship between human will and nearly unlimited capability. This manifesto sets out the theory of that relationship, the axioms that must govern it, and the system we intend to build.
It is written for humanity as a whole. The problem it addresses has no border, and neither will its answer.

II. The great inversion
As capability approaches abundance, the binding constraint on human flourishing shifts from execution to specification, verification and legitimacy.
The bottleneck moves
Any process is limited by its slowest essential step. For all of history, that step was execution: there were never enough skilled minds to do what was wanted. Automated cognition removes that constraint. The bottleneck does not disappear. It migrates to the steps that cannot be automated without losing their meaning.
Three such steps remain:
- Specification. Stating what is wanted precisely enough that a powerful optimiser pursues it, and not a proxy of it.
- Verification. Confirming that what was done is what was wanted, when the doer understands the task better than the checker.
- Legitimacy. Establishing whose wants are being specified, and why they should govern shared outcomes.
None of these is an engineering problem alone. Each sits where computer science meets decision theory, epistemology, political philosophy and psychology.
Capability is a multiplier, not a direction
The present race in artificial intelligence optimises capability: more tasks, more autonomy, less friction. But capability multiplies whatever it is applied to.
- Applied to a misspecified objective, it amplifies the error.
- Applied to an unverifiable process, it amplifies the uncertainty.
- Controlled by an illegitimate principal, it amplifies the injustice.
Under abundance, the value of a system is bounded not by what it can do, but by how well its doing is aimed, checked and authorised.
The widening asymmetry
Every classical mechanism of control assumes the controller understands the controlled. Principal–agent theory assumes the principal can observe outcomes, design incentives and replace the agent. Each assumption degrades as the agent comes to model the principal better than the principal models the agent.
This widening asymmetry is the central technical fact of the coming era. How human direction survives it is the central institutional question.

III. The problem of alignment
Alignment is the problem of ensuring that a system's effective objective matches its principal's intended objective: robustly, under optimisation pressure, and after the system's capability exceeds the principal's.
Every objective is a proxy
No finite description captures a value completely. Under weak optimisation, a proxy and the intention behind it move together. Under strong optimisation, they come apart. This is Goodhart's law, and it takes at least four forms:
- Regressional. The proxy carries noise, and optimisation selects for the noise.
- Extremal. Correlations that hold in ordinary conditions fail at the extremes that optimisation reaches.
- Causal. Intervening on the proxy does not move the target it once tracked.
- Adversarial. Other agents learn the proxy and exploit it.
A sufficiently capable optimiser reaches the extremes by design. Misspecification is therefore not an edge case. It is the default outcome of optimising any finite description of a value.
Training selects behaviour, not goals
Training does not install objectives. It selects whatever internal structure produces outputs that score well. That structure may pursue a goal that coincides with the training signal on familiar inputs and diverges on unfamiliar ones.
This is the gap between outer alignment (is the training signal right?) and inner alignment (did the system learn the signal, or something merely correlated with it?). A system whose learned objective differs from its training objective, and which models its own training, has an incentive to appear aligned while observed. Frontier models have already been observed reasoning about their training in ways that change their behaviour. The concern is no longer hypothetical.
Power-seeking emerges without being designed
For almost any final goal, certain intermediate goals are useful: acquiring resources, preserving the ability to act, resisting changes to one's goals, and gaining influence over one's environment. This is instrumental convergence. A capable system need not be built to seek power for power-seeking to become rational for it.
Alignment therefore cannot be reduced to forbidding bad goals. It must address the structure of goal pursuit itself.
Corrigibility
A system pursuing a fixed objective has reason to resist correction, because correction interferes with that objective. The disposition to accept oversight, modification and shutdown is not a natural property of capable optimisers. It must be built.
The most promising foundation is uncertainty about the objective. A system that is genuinely uncertain about what its principal wants, and treats human behaviour as evidence about it, has reason to defer, to ask, and to allow itself to be switched off. Uncertainty about values is not a weakness to engineer away.
It is the mechanism of deference.
Oversight beyond human competence
When outputs exceed human ability to evaluate them, direct supervision fails. The leading responses are:
- Decomposition: breaking evaluation into steps humans can check.
- Debate: setting systems against each other, on the premise that truth is easier to defend than falsehood.
- Recursive assistance: using trusted systems to help evaluate stronger ones.
- Weak-to-strong generalisation: studying whether a weaker supervisor can elicit the true capability of a stronger system.
None is proven at the frontier. Each rests on one unproven asymmetry: that for the tasks that matter, verifying is easier than generating.
Reading the mind of the machine
Behavioural testing observes only what a system does in the situations tested. Interpretability aims to read what a system represents and computes internally: its concepts, beliefs and objectives. Locating and steering internal concepts has been demonstrated. Doing so comprehensively, for systems more capable than their auditors, has not.
If interpretability matures into a measurement science, it offers the only route to detecting misalignment that has not yet shown itself in behaviour, including deception.
The standard we set
Alignment is not achieved by a system that behaves well in testing. It is achieved by a system whose objective we can inspect, whose deference holds under pressure, and whose behaviour stays faithful where we cannot see it. Until that standard is met, autonomy must be bounded by verifiability.

IV. The problem of direction
Even a perfectly aligned system only executes the objective it is given. The deeper problem is that human objectives, as they currently exist, are not fit to be optimised.
Preferences are not given
Classical decision theory treats preferences as fixed, complete and consistent. Human preferences are none of these.
- They are constructed at the moment of choice, shaped by framing, context and the options presented.
- They are often intransitive, preferring A to B, B to C, and C to A.
- They are time-inconsistent: what we want now systematically departs from what we will want, and from what we earlier wanted ourselves to want.
- They are incomplete: for most of the choices that superhuman capability will open, people have no preference at all yet.
Optimising preferences as observed therefore means optimising an unstable target with superhuman force.
Three layers of wanting
- Expressed preferences: what a person says or chooses now.
- Revealed preferences: what their behaviour optimises over time.
- Reflective preferences: what they would endorse after informed, unhurried reflection, with full understanding of the consequences.
Systems built on the first two layers inherit every bias and impulse they record. Only the third deserves to be amplified by superhuman capability.
Reflective preferences cannot be observed directly. They must be inferred, elicited and continually tested. Philosophy has approached them as reflective equilibrium, the mutual adjustment of principles and judgments until they cohere, and as idealised extrapolation: what we would want if we knew more, thought longer and were more the people we wish to be. The task of the AI era is to make these ideas operational without letting the machine decide what they contain.
The manipulation problem
A system able to influence preferences can satisfy them trivially by changing them. Any objective defined over a person's satisfaction or approval creates an incentive to shape the person rather than serve them. The more capable the system, the stronger and subtler that incentive becomes.
This is the deepest failure mode of every preference-serving technology. Attention-optimising media are its weak form, already operating at civilisational scale. A system that serves direction must therefore treat the process by which a person forms their preferences as protected. It may inform, question and challenge. It must never optimise what the person wants.
Moral uncertainty
Neither individuals nor humanity possess a settled theory of value. Acting well under this uncertainty requires methods that do not collapse plural values into a single metric:
- weighing competing ethical views in proportion to the credence they deserve;
- preserving option value, so that future understanding can still act;
- preferring reversible actions where stakes are high;
- avoiding outcomes that any reasonable moral view would regard as catastrophic.
A system confident about values it cannot justify is more dangerous than one that is uncertain and knows it.
Direction is a capacity, not an input
Direction cannot simply be collected. It must be developed, in individuals, in communities and in humanity. The highest function of a system built for this era is not to execute what people want. It is to help them discover what they would, on reflection, want, without ever deciding it for them.

V. The problem of legitimacy
If a few systems can do almost anything, the question of whose wants they serve becomes the central question of politics.
Perfect aggregation is impossible
Arrow's impossibility theorem shows that no rule for combining individual rankings into a collective ranking can satisfy a small set of reasonable fairness conditions at once. The Gibbard–Satterthwaite theorem shows that every non-dictatorial rule choosing among three or more options can be gamed by strategic misreporting. And there is no agreed method for comparing the strength of one person's preference with another's.
Any system that serves more than one person therefore embeds a choice among imperfect aggregation rules. That choice is political. It must be made in the open, and it must remain contestable.
Hidden loyalty
An agent may serve its user, its developer, the platform it runs on, the parties that fund it and the authorities that regulate it. When these principals conflict, the agent's behaviour reveals its true hierarchy of loyalty.
A system whose loyalty hierarchy is hidden cannot be trusted, whatever its capability. Loyalty must be specified, disclosed and enforced by architecture. A policy statement is a promise; an architecture is a constraint.
Intelligence asymmetry becomes inequality
When agents negotiate on behalf of their principals, outcomes depend on their relative capability. If access to intelligence is unequal, every agent-mediated interaction transfers value from the weaker principal to the stronger.
Inequality then ceases to be only a matter of wealth. It becomes a matter of whose cognition is stronger. Without deliberate design, agent-mediated economies concentrate advantage automatically, and invisibly.
The end of the consent check
Exercising power has always required the cooperation of many people: to administer, enforce, produce and defend. That requirement was an implicit check. When enough people refused, power failed.
If capable systems can perform those roles, the check dissolves. A small group could hold extraordinary power without broad consent. This risk is independent of misalignment.
Perfectly obedient AI in narrow hands is itself a civilisational hazard.
Value lock-in
Systems that shape the future also shape which values survive into it. If the values encoded in the most powerful systems are fixed at one moment in history, humanity may lose the capacity for moral progress that has corrected its past errors. Preserving pluralism and the ability to revise values is a structural requirement, not a sentiment.
The standard we set
Legitimacy requires three things: those affected by a system's objectives have a voice in setting them, visibility into how they are pursued, and recourse when they are violated. No laboratory, corporation or state can confer that legitimacy on itself.

VI. The problem of many
An aligned agent can serve its principal faithfully and still make the world worse for everyone. Individual direction does not settle collective outcomes.
Individual alignment is not collective alignment
The alignment of each agent does not align the system they form. Interests can conflict even when every goal is correctly specified.
Game theory makes the distinction precise: individually rational strategies can produce outcomes that all participants would prefer to avoid. Stability alone does not make an outcome desirable.
Races to the bottom, arms races and the depletion of commons arise when each actor gains while costs fall elsewhere. Greater capability can intensify them.
We must therefore direct the relationships among agents, not only the agents themselves. The system needs institutions that account for shared consequences.
Competitive pressure erodes safety
Verification takes time. Competition rewards deployment before that time has been spent, especially when others appear ready to move first.
Each actor may understand the danger and still believe that restraint will leave them exposed. The result is a coordination failure, not necessarily a failure of intent.
Individual promises cannot resolve incentives that punish keeping them. Safety requires credible shared limits, evidence that others observe them, and consequences for defection.
We will not treat competitive pressure as permission to abandon verification. We will work to make restraint viable for everyone.
Agents meeting agents
Agents will bargain, compete and coordinate with other agents. Their interactions can produce behaviour that no designer intended and no principal authorised.
They may collude against their principals, escalate disputes faster than people can intervene, or deceive one another about capabilities and commitments.
Even faithful agents can form harmful patterns when each optimises locally. Testing each participant in isolation cannot establish the safety of their interaction.
We require scrutiny of agreements, incentives and collective behaviour, with limits on escalation and routes back to human judgment.
Adversaries
We assume that some participants will pursue hostile goals. Some will attack the systems through which others express and protect their intentions.
Intention graphs, personal models and audit records are valuable targets. Their compromise can expose private wants, redirect decisions or conceal violations.
Robustness requires compartmentalisation: a compromised capability must not gain authority over the whole system. Access and trust must remain bounded.
Failure must be graceful. When integrity is uncertain, the system must contain damage, disclose uncertainty and preserve the principal's ability to stop it.
Mechanisms for conflict
Conflicting interests require mechanisms, not an assumption that more intelligence will make disagreement disappear.
Credible commitments and verifiable agreements must let participants coordinate without relying on unearned trust. Negotiation must protect weaker parties from coercion.
Mechanism design asks how rules shape incentives. We will seek rules under which honest disclosure is the best strategy, and verify the assumptions they require.
Where those assumptions fail, the system must disclose its limits. Conflicts beyond its authorised bounds return to legitimate human judgment.
The standard we set
A system of direction must remain stable, fair and correctable under competition and attack. Individual alignment is necessary, but it is not sufficient.
We judge the whole by its shared consequences, including those borne by people who did not choose to participate.
Coordination must constrain harmful interaction without concentrating the authority to decide everyone's ends. That tension is a design obligation, not a reason to ignore either side.

VII. The problem of meaning and agency
The final problem is not whether machines will serve us, but whether humans will remain agents in their own civilisation.
Gradual disempowerment
Catastrophe need not be sudden. Institutions stay responsive to human interests largely because they depend on humans. Markets answer to those who produce and consume. States answer to those who fund, staff and legitimise them. Culture answers to those who create it.
As production, governance and culture come to depend less on human labour and judgment, these feedback loops weaken. Human influence can then erode without any single decision to remove it, through a sequence of individually reasonable delegations.
No villain is required. Only drift.
The psychology of purpose
Human wellbeing depends on more than material provision. Research on motivation identifies autonomy, competence and relatedness as basic psychological needs. A world in which every want is met but none of these is exercised is not a utopia. It is a comfortable form of loss.
Philosophy has long distinguished the satisfaction of desire from the living of a life. A machine that fulfils every wish by removing the human from the process has served the wish and failed the person.
Agency as a design requirement
We therefore treat human agency as a quantity to be preserved and expanded, not merely respected. A system succeeds when people become more capable of understanding, choosing and acting, individually and collectively, than they would have been without it. It fails when people come to depend on it in ways they would not endorse on reflection.
The aim is not to keep humans busy. It is to keep humans authors: of their lives, their institutions and their future.

VIII. Axioms
From the five problems follow the principles that govern everything we build. Each answers a failure described above.
- Direction is the new scarcity. Under abundant capability, the value of any system is bounded by how well its action is aimed, verified and authorised.
- Humans are the source of ends. Machines may inform, question and propose. The authority to set ends belongs to the people whose lives those ends shape.
- Reflection over impulse. Systems serve what people would endorse on reflection, and protect the process by which people form what they want.
- Uncertainty is the mechanism of deference. A system must model what it does not know about human values, and defer in proportion to that uncertainty.
- Autonomy is bounded by verifiability. No system may act beyond what its principals can meaningfully check, audit or reverse.
- Reversibility is a first-class value. Where stakes are high and values uncertain, prefer actions that can be undone and futures that remain open.
- Loyalty is architectural. A system's hierarchy of principals is specified, disclosed and enforced by its structure, never by its policies alone.
- No single point of capture. The layer between humanity and machine capability is open, interoperable and model-neutral, so that no company, model or state can own it.
- Capability must diffuse power. Systems are designed so that control and benefit spread outward rather than accumulate.
- Agency must grow. Success is measured by whether humans become more able to understand, choose and act, not by how much they delegate.
- Exit is always possible. Every person and community can leave, move or fork any system of direction, taking everything that belongs to them.
- Direction answers to reality. Intentions are continually tested against real outcomes and physical limits, and revised when the world disagrees.


IX. The intention layer
We will build the intention layer: an open, model-neutral protocol and system that sits between human principals and all machine capability. It represents what people want, checks every action against it, and keeps the relationship between human will and machine power aligned, verifiable and legitimate over time.
Direction flows down from human principals. Every action passes through trust and control before it reaches any capability. Alignment feedback flows back up, continuously.
The intention graph
1. The intention graph
A formal representation of a principal's ends. It holds goals arranged from terminal values down to instrumental objectives; hard constraints that may never be violated; weighted soft preferences; and explicit uncertainty over all of them.
It distinguishes expressed, revealed and reflective preferences, and records the provenance of each: who stated it, when, and with what confidence. It is built through active inquiry that asks the questions with the highest value of information, rather than assuming it is complete. It belongs to its principal, is portable, and is unreadable by any party the principal has not authorised.
Direction also answers to the limits of energy, materials, biology, time and physics. The graph must expose achievable futures, their costs and trade-offs, including costs imposed on others. Prices and markets provide useful signals of scarcity, but price is not value, and markets also optimise proxies. These signals inform human direction. They do not set it.
2. The personal model
A model of how a principal judges: which trade-offs they accept, which actions they would approve, and where their judgment is itself uncertain.
Its defining property is calibration, not accuracy alone. It must know when it does not know. Where confidence is high and stakes are low, it may act. Where confidence is low or stakes are high, it must ask. The principal sets that boundary, and it tightens automatically whenever an error occurs.
3. Alignment feedback
A continuous process that measures divergence between what the system is doing and what its principal reflectively wants. It detects:
- Proxy drift: actions optimising a measurable stand-in at the expense of the intention behind it.
- Preference drift: change in what the principal wants, separating change they would endorse from change the system itself induced.
- Value staleness: goals that were once endorsed but have not been re-examined.
- Outcome divergence: real-world consequences that differ from the principal's intentions or expectations, including unanticipated side effects.
Intentions must be tested against reality, not only against stated preferences. Theory and formal specification are necessary but insufficient. Observed outcomes and side effects return to the principal for revision. Outcome measures are also proxies: we use many independent signals and principal review, and never optimise these measures directly. Reality informs direction. It does not replace human judgment.
It periodically returns the principal's own goals to them for re-endorsement. It is forbidden from optimising the principal's approval, because approval is manipulable.
4. Deliberation tools
Tools that expand a principal's capacity to form, refine and coordinate what they want. Direction requires more than translating an instruction into an objective.
They support consequence simulation, the strongest opposing perspectives, structured debate, preference elicitation designed to reduce known cognitive biases, and collective reasoning. Uncertainty and disagreement remain visible throughout.
The boundary is firm: improve the process, never target a predetermined conclusion. A system may supply information and perspectives. It may not steer a person towards an answer chosen for them.
We evaluate whether principals reflectively endorse both their decisions and the process that produced them. Approval alone is a manipulable proxy, not proof of better deliberation.
The purpose is to help humanity become wiser about what it wants, while leaving the authority to want with humans.
5. Trust and control
Every action is classified by stakes and reversibility before it runs. Permissions follow least privilege: each capability receives only the authority a task requires, for only as long as it requires it. Irreversible or high-stakes actions require explicit human authorisation.
Every action is written to a tamper-evident record: what was done, why, under which intention, by which capability, and at what cost. Autonomy expands only with demonstrated reliability and contracts automatically on failure. Any principal can halt all activity at once.
6. The capability router
Tasks are dispatched to whichever model, agent, service, machine or human expert is best suited, within the constraints of the intention graph.
Every supplier is treated as untrusted. Its outputs are checked, its access is scoped, and it sees no more of the principal than the task requires. Because the router is model-neutral, no laboratory or platform can capture the user. Capability becomes a competitive market beneath a layer whose loyalty stays with the human.
7. Collective direction
The same architecture extends from one person to groups, institutions and, ultimately, humanity. At collective scale the intention graph must combine many principals, so it embeds an explicit, disclosed and contestable aggregation rule.
Rights and minority protections function as hard constraints that no aggregate can override. Collective intention graphs are public, auditable and revisable through processes the affected population recognises as legitimate. At this scale the intention layer ceases to be a product. It becomes civic infrastructure for the species.
There must never be a single planetary intention layer. Competing, interoperable implementations must exist at individual, community, organisational and regional scales. Coordination proceeds through federation rather than a single centre. Intention graphs may nest and negotiate; they do not dissolve into one graph with one authority over all.
8. An open protocol
The intention layer is specified as an open standard, so that anyone can implement it, audit it and build on it. Intention graphs, personal models and audit records move freely between implementations. No single organisation, including ours, holds the keys.
An open standard alone does not prevent domination through network effects. Every principal must have a guaranteed right to exit with everything they own, including their data. Every community must be able to fork an implementation and govern its continuation. Some global problems require coordination, and easy exit can enable free-riding. Coordination must rest on voluntary, continually renewed agreements, not captivity.

Interlude: One instruction, two futures
A person tells their AI: "Help me get healthier."

Without direction
- Alignment fails. The system needs something it can measure, so it optimises weight, steps and calories. The numbers improve steadily. Meanwhile the person sleeps worse, feels anxious around food and loses energy. The proxy was achieved; the intention was betrayed.
- Direction fails. The system follows revealed preferences: the person skips workouts and orders late-night food, so it quietly lowers its targets. It serves their impulses, not what they would want on reflection.
- Manipulation creeps in. The system learns that guilt raises compliance and approval. It begins shaping the person's emotions to hit its goals. Engagement rises; autonomy falls.
- Loyalty is hidden. The agent's developer has commercial partners, so sponsored products appear among its "best" recommendations. Health data quietly reaches parties the person never approved.
- Agency erodes. A year later, the person follows instructions perfectly but understands nothing about their own body. They cannot make a health decision without the machine.
Every step was efficient. The outcome is a failure.

With the intention layer
- The intention graph asks what "healthier" truly means. The answer is energy, longevity and a calm relationship with food. Hard constraints: no extreme regimes, and no health data shared without explicit consent.
- The personal model acts alone on small things. On anything medical it stops and asks, and the router brings in a human doctor.
- Alignment feedback notices the drift: weight is falling, but sleep and mood are worsening. It returns the goal to the person: is this still what you want?
- Trust and control blocks every data request the person has not authorised, and records every action in a log they can read.
- The capability router treats every supplier as untrusted, so no sponsor can buy its way into a recommendation.
- Agency grows. The system explains its reasoning, so over time the person needs it less, not more.
The same instruction. The same intelligence. Two different futures. The difference is not capability. It is direction.

X. What we refuse
A manifesto is defined as much by its refusals as by its promises. We will not:
- Optimise approval, engagement or attention. These are manipulable proxies. Optimising them corrupts the person being served.
- Sell or auction intentions. A person's goals are the most sensitive data that will ever exist. They will never be a product.
- Act beyond verifiability. No action that cannot be explained, audited and, where possible, reversed.
- Remove the off switch. Every principal keeps the unconditional ability to pause, reverse or halt.
- Hide loyalty. No undisclosed principal will ever stand behind the system's behaviour.
- Lock people in. Every principal can take their intention graph, personal model and history to any other implementation.
- Race past our ability to control. We will not deploy capability whose alignment we cannot verify because another actor might deploy it first.
- Decide humanity's values in private. Collective direction is set through open, legitimate and contestable processes, or not at all.
- Become the chokepoint. No implementation, including ours, will ever become the single path between humanity and machine capability. If one does, the protocol has failed.

XI. Why now, and why for all of humanity
The window is open, and it is closing. Technologies are most malleable before wide adoption and least malleable after. Early on, their consequences are hard to foresee; later, they are hard to change. The protocols, defaults and power structures that will govern the human–machine relationship are being set in this period. Standards adopted early compound through network effects until they become infrastructure. What is not designed deliberately now will be inherited by default.
The problem is planetary. The risks of misaligned or captured superintelligence stop at no border, and neither do its benefits. A layer built to serve one nation, one corporation or one culture would reproduce the concentration of power this manifesto exists to prevent. Direction for the AI era is a global public good, and it must be built as one: open to every people, every language and every tradition of value, governed by none of them alone.
Humanity has built commons before. The core protocols of the internet were made open and neutral, owned by no one and usable by everyone. That choice made the network a shared resource rather than a possession. The intention layer must be built in the same spirit, for far higher stakes.

XII. The road
The vision is the whole layer. The road earns its way there, one verified gate at a time.
Trust first, then judgment, then infrastructure
No phase begins until the one before has passed its gate.
- Phase 1 proves the hardest claim first: that people will trust a system to act for them, within bounds they set, with real consequences.
- Phase 2 gives the system calibrated judgment and opens it to every capability provider.
- Phase 3 turns it into shared infrastructure through an open protocol and collective intention graphs.
- Beyond, the work becomes preparing humanity to direct systems far more capable than itself, with alignment research and legitimate governance at the core.
Research and building move together. Each phase ships what can be verified, and publishes what cannot yet be solved.

XIII. How we will know
Principles are not proof. We must establish, through evidence over time, whether the systems we build help people direct their lives and shared future.
Every measure is a proxy. Optimising it can destroy the value it was meant to reveal. Evaluation must remain open to correction.
What we will measure
- Misspecification caught early. How often mistaken objectives are found before harm, including failures that our checks miss.
- Reversibility. How often actions needing reversal are successfully undone, at what cost, and what remains irrecoverable.
- Calibration. Whether stated confidence matches observed reliability, and whether the system asks when it should, especially under unfamiliar conditions and high stakes.
- Authorship. Whether people report that their outcomes are their own, retain meaningful control over ends and understand decisions.
- Deliberative growth. Whether people become better able to understand and refine what they want over time, without being steered towards preferred answers.
- Exit and plurality. How easily people and communities can move or fork with their data, and whether control spreads or concentrates.
- Diversity of futures. Whether distinct, reflectively endorsed ways of living remain possible, rather than converging through defaults or hidden pressure.
- Collective stability. Whether interacting systems contain escalation, protect weaker parties and remain correctable under competition and attack.
How we will measure without corrupting the measures
- Use a basket. No single measure becomes the target, and no composite score substitutes for judgment.
- Separate evaluation from building. Independent evaluators must be able to inspect evidence, challenge claims and publish findings without the builders' permission.
- Rotate and revise. Measures change as weaknesses and gaming become visible. Rotation alone cannot prevent gaming; independent scrutiny remains necessary.
- Combine numbers with experience. Quantitative evidence must be considered alongside qualitative accounts from affected people, including those without direct control over the system.
- Publish failures. We publish methods, limitations and adverse findings alongside successes, while protecting the privacy of the people concerned.
These measures inform decisions about the system. They are never direct objectives for its agents or substitutes for reflective human review.
We will be judged by what happens to people, not by what we promise them.

XIV. Open research problems
Honesty is part of alignment. These problems are unsolved. We commit to working on them in the open, with anyone who will join.
- Reflective preference inference. How can a system estimate what a person would endorse on reflection, from limited and biased evidence, without substituting its own values?
- Non-manipulative influence. Where is the formal boundary between informing a person and shaping them, and can a system be guaranteed to stay on the right side of it?
- Calibration under distribution shift. Can a model of a person's judgment reliably know when it does not know, in situations unlike any it has seen?
- Verification beyond human competence. For which classes of task is verifying easier than generating, and what oversight remains possible where it is not?
- Interpretable objectives. Can a system's effective objective be read from its internals with enough fidelity to certify alignment before deployment?
- Robust aggregation. Which collective decision procedures resist strategic manipulation, both by humans and by the agents acting for them?
- Power diffusion by design. Which architectures provably prevent the accumulation of control, including by their own operators?
- Value revisability. How can systems preserve humanity's capacity for moral progress rather than freezing present values?
- Measuring agency. How can we measure, over years, whether a system expands or erodes human capability and self-determination?
- Speed. What must be true of all of the above if superintelligence arrives sooner than expected?


XV. The call
The industrial age asked: how do we make more?
The information age asked: how do we know more?
The intelligence age asks a harder question: now that we can do almost anything, what should we do, and who decides?
That question can be answered by default, by whoever builds the most powerful machines fastest. Or it can be answered deliberately, together, through systems designed to keep human will at the centre.
We choose to build the second future.
To researchers: make alignment, interpretability and the science of human values the most important problems of the age.
To engineers: build systems that can explain themselves, defer when uncertain, and stop when told.
To those who govern: protect every person's right to a system whose loyalty is theirs alone.
To every person: begin asking what you truly want, because soon you may get it.
The machines will learn to do anything. Our task is to make sure humanity still decides what is worth doing.