← All Paramerge white papers

Paramerge White Paper

Pricing Consensus

A Thought Experiment in Exit-Aware Collective Decisions Using Option Chains

Stephen Lieberman

Paramerge, Real-World AI Governance Center

August 2026, Draft 1.0, revised after a second adversarial review

stephen@paramerge.com · paramerge.com

Abstract

Current AI-mediated deliberation systems find consensus by generating candidate statements and selecting among them with a preference-aggregation rule. That design assumes participants cannot leave. In practice they can, and they do, and agreements endorsed under a forced choice can fail once people are free to pursue a smaller coalition that fits them better. This specification describes a different approach. Each candidate policy is treated as an underlying, each setting of its main parameter as a strike, and each participant's declared willingness to contribute at each strike as a bid. The aggregate of bids across strikes is a priced option chain. The mathematics of the mechanism are constructions inspired by classical results, not applications of them, and the paper says exactly which hypotheses of each theorem hold here and which fail. A total-benefit test inspired by the public goods literature tells the group where along the chain aggregate willingness meets cost. A core test in the spirit of Foley and of participatory budgeting asks whether any subgroup would rather leave, and answers one-sidedly. A curvature reading inspired by Breeden and Litzenberger treats the shape of the chain as a diagnostic of where declared contributions concentrate and where they divide, and earlier drafts' claim that this reading recovers the group's implied probability distribution over outcomes is withdrawn in favor of a conditional statement whose failure conditions are enumerated in the text. Exit is a parameter of every participant's input rather than an afterthought, disagreement is visible in the individual schedules and in the shape of the chain, and every output is traceable to a logged input and a stated rule. The mechanism's nearest rigorous neighbors price beliefs, as scoring-rule markets do, price a single allocation, as the Lindahl and core mechanisms do, or price the right to abandon a commitment, as leveled commitment contracts do; no published mechanism known to the author prices collective commitment as a chain across the settings of a policy, and the novelty claimed is confined to that structure, the diagnostics it yields, and the exit-durability standard proposed to validate it. The document is a thought experiment. It specifies the objects, inputs, aggregation, decision rule, stability test, sensitivity outputs, scope conditions, and assumptions at the level needed to build and test the mechanism in silico, states plainly what the method is not, and records in an appendix two rounds of adversarial review, ten expert perspectives each, and how every objection was resolved or why it stands.

1. The problem with forced-choice consensus

The leading AI deliberation system, the Habermas Machine, generates candidate group statements, predicts each participant's ranking of them with a personalized reward model, and selects a winner by simulating an election under the Schulze rule (Tessler et al., 2024; Schulze, 2011). Its predecessor selected statements by maximizing a social welfare function over predicted approval (Bakker et al., 2022). Generative social choice constructs a proportionally representative slate of statements under stated representation guarantees, a design built to seat minority positions rather than to crown a single most-endorsed item (Fish et al., 2024). Bridging systems such as Community Notes surface content approved across an opinion divide. These are different rules with different virtues, and they share one structural assumption. The participant set is fixed, no input describes what any participant would do instead, and the rule selects on the premise that everyone remains in the room.

That assumption sets aside a variable that matters for whether an agreement lasts. Participants have outside options. Hirschman (1970) named the choice between exit and voice and showed that the availability of exit changes the meaning of every act of voice. Bargaining theory formalized it. An outside option that exceeds a party's negotiated share becomes a binding constraint, and any agreement that ignores it will be renegotiated or abandoned (Osborne and Rubinstein, 1990). The prediction is sharp, an outside option affects the bargain only when it binds, and when it binds it sets the terms, and it has supportive experimental evidence against the conventional alternative (Binmore, Shaked, and Sutton, 1989); when both sides hold outside options the set of equilibria widens to include delay (Ponsatí and Sákovics, 1998). Cooperative game theory generalized it to groups. An outcome outside the core admits a blocking coalition by definition, and core-stable outcomes need not exist at all (Bogomolnaia and Jackson, 2002). And exit does not only threaten agreements; it builds alternatives. Tiebout (1956) and Buchanan (1965) showed that when leaving is cheap, people sort into jurisdictions and clubs that fit them rather than compromise where they stand, which is coalition formation by exit under another name, and which in Tiebout's model rests on assumed full mobility and full information, assumptions that distribute like wealth.

The honest state of the evidence deserves one paragraph, because two earlier drafts overstated it. In the largest deployed bridging system, 30.2 percent of notes that reach displayed consensus later lose their helpful status (Chuai, Lenzini, and Pröllochs, 2026). The measured mechanism is not exit. Display triggers a surge of re-rating, contributors with viewpoints similar to the note's author increase supportive ratings while dissimilar contributors increase negative ones, and the resulting post-display polarization drives much of the disappearance. That is evidence that bridged consensus is unstable once displayed, under continued and arriving voice, not evidence that exit destroys it, and it carries a second lesson this specification takes personally in Section 11, that displaying an aggregate changes the behavior that produced it. The Habermas Machine study itself reports that initial group statements weighted minority opinions in proportion to their empirical share and that statements revised after critique rounds weighted them above it, a shift the study's authors read as the mediator incorporating dissent while respecting the majority position (Tessler et al., 2024). The point taken here is narrower than earlier drafts made it, that every aggregation rule has weighting properties the participants did not choose and cannot see, whichever direction they run. On manipulation, learned manipulability under limited information varies sharply by rule, Borda is highly manipulable, Instant Runoff resists despite full-information vulnerability, and the Condorcet methods tested were the least manipulable of the eight studied; Schulze itself, a Condorcet method, was not tested (Holliday, Kristoffersen, and Pacuit, 2025). Manipulability is therefore not the indictment this paper brings. The indictment is missing information. None of these systems takes a declared outside option as an input, none models coalition formation, and no field measurement of exit from a deployed deliberation system appears to exist. The motivating premise of this paper stands on Hirschman, on bargaining theory, and on coalition theory, and producing the missing measurement is part of the validation plan in Section 14.

2. The proposal in one paragraph

Treat each candidate policy as an investment. Let each participant state, for each setting of the policy's main parameter, what they would contribute from a shared resource to enact the policy at that setting. Aggregate those contributions. The result is a priced option chain for that policy, one price per strike. Price every candidate policy the same way. Then compare the chains, and the individual schedules beneath them, to decide where the group should commit its resource. Because every participant's contribution schedule embeds a level below which they contribute nothing, exit is built into the inputs. Because the chain is a curve rather than a point, its shape shows where the group agrees, where it hesitates, and where it is silently divided. Because the prices are declared inputs passed through stated rules, every output can be traced to a logged input and the rule that transformed it. This is offered as a thought experiment, meaning a fully specified design whose claims are testable, not a report of a deployed system.

3. The neighbors and the white space

The idea of using prices to make collective decisions has a rigorous history, and this specification should be read against it rather than beside it. Hanson's logarithmic market scoring rule (Hanson, 2003) is the canonical mechanism for aggregating dispersed beliefs into a price, a subsidized automated market maker with bounded worst-case loss that is incentive compatible for myopic risk-neutral traders. Futarchy (Hanson, 2013) builds governance on it, voting on values and betting on beliefs. Both are belief machines, and it must be said plainly that a scoring-rule market over a continuous outcome already yields a full group-implied probability distribution as its native output, disciplined by settlement against the realized outcome. What such a market does not price is commitment. It needs an observable event to settle against, and it models no one's option to leave.

The exit constraint has a rigorous home of its own. Lindahl (1919) proposed pricing a public good through personalized prices, and Foley (1970) proved that a Lindahl equilibrium allocation lies in the core. Equilibrium is the load-bearing word, and Section 6 returns to it, because this mechanism computes no equilibrium and inherits no core property. Fain, Goel, and Munagala (2016) brought the core to participatory budgeting, where the Lindahl equilibrium they characterize is always a core solution for the utility class they study, and where they also develop a randomized approximately dominant-strategy truthful selection rule. Garg, Goel, and Plaut (2021) built markets for public decision-making on the same foundations. Budish (2011) disciplined a market for indivisible goods with artificial budgets that are nearly equal, unequal but arbitrarily close together, because exact equality breaks existence, and his mechanism is strategyproof in the large-market limit. The nearest neighbor of all is a 2026 preprint by Song and Nguyen (arXiv:2603.04312, v1, 4 March 2026), which extends Lindahl equilibrium to voters with ordinal preferences over a finite outcome set when monetary transfers are not allowed, and proves approximate core guarantees, including a deterministic factor of 6.24 that improves a previous bound of 32. Its constraint is coalitional, a core condition on what a group's share of the electorate can support, not an individual outside-option constraint, and its motivation runs opposite to this specification's, since it removes money from the voting problem precisely because willingness to pay is a contested weight there. A cardinal, numeraire-denominated design owes that argument an answer rather than a citation. The answer this specification gives is the equal-budget artificial numeraire as the default in Section 6, and the scope conditions in Section 13 that mark where no numeraire is defensible.

Two more neighbors complete the honest map. Multi-agent systems research prices exit directly. Leveled commitment contracts make every commitment breachable at a stated decommitment penalty and analyze strategic breach by self-interested agents (Sandholm and Lesser, 2001). That is a priced right to abandon an agreement, with no chain across settings and no diagnostics, and any paper claiming that existing mechanisms do not price exit must name it. And the nearest published use of option pricing on a political object is a Black and Scholes valuation of a two-party pre-electoral coalition agreement (Mitrović, 2024), which models party support as a stochastic process and solves a single valuation against a genuine random underlying. It is a valuation exercise rather than an aggregation mechanism, it builds no chain and recovers no distribution, and unlike this specification it has a random variable, a difference Section 8 will make much of.

The white space this specification claims is therefore deliberately narrow. Exit as a constraint is standard mechanism design under the name individual rationality, and no novelty is claimed for it. Distribution-shaped outputs are the native product of scoring-rule markets, and no novelty is claimed for those either. What appears to be unpublished is the combination this paper specifies. The decision is priced as a chain of binding commitments across every setting of a policy rather than at one point, so the curve itself is the deliverable; the diagnostics read disagreement off the declared schedules and the shape of that curve; and the validation standard is durability under exit in a grounded testbed, a standard to which no aggregation mechanism in the deliberation literature has yet been held. Social choice also has instruments of its own for non-participation, abstention rules and veto cores among them, and the specification does not claim to supersede them. It claims a different output, a priced and interpretable curve rather than a selected winner.

The time dimension of the mechanism has a rigorous home as well. Real options theory (McDonald and Siegel, 1986; Dixit and Pindyck, 1994) prices the option value of waiting to make an irreversible commitment under uncertainty, and its central lesson is that a holder of a live option to wait does not rationally commit at break-even. A group deciding under a validity horizon holds exactly that position. Sections 7 and 9 take both halves of the lesson, the value of waiting and the fact that a break-even rule ignores it.

4. Objects

Underlying. A candidate policy or coherent policy set. Each is priced separately. Qualitatively different policies are different underlyings and are not mixed within a chain.

Strike. A setting of the underlying's main parameter. Stringency of a threshold, level of coverage, size of a budget cap, intensity of enforcement. A chain is the underlying evaluated across an ordered set of strikes. The choice of the strike axis is the most consequential act in the procedure, it is an agenda decision, and Section 11 makes the axis a logged, contestable input rather than a given.

Numeraire. The single common thing the group spends to enact a policy. Budget, staff hours, credits, or political capital. It may be real or artificial, but it must be one thing and it must be declared before inputs are collected. All prices are denominated in it. The recommended default is an artificial numeraire with equal budgets, for reasons Section 6 states; the real-numeraire path changes what the mechanism measures and is flagged wherever it does.

Outside option. A named alternative available to a participant if the policy is not enacted, including a smaller coalition pursuing a different underlying. Outside options may be declared by participants or proposed by the epistemic mapping layer described in the companion white paper, subject to the provenance, ratification, and contestation rules of Sections 11 and 12.

Outcome variable. Optionally, a single named quantity, in stated units with a stated measurement, by which the group agrees the policy's result will be judged. Declaring one is what licenses the conditional distribution reading in Section 8. Without it, Section 8 reports curvature and nothing more.

Decision horizon. The date by which the decision must be made. It defines the chain's term.

Validity horizon. The date past which the policy as currently understood ceases to be valid, because the conditions, actors, or resources it assumes will have changed enough that the participants' valuations no longer apply. Past the validity horizon the chain is void and must be re-priced from fresh inputs. This is the expiry of the option, and the value of waiting that it extinguishes is the option value priced by the real options literature cited in Section 3.

Safety, fairness, equity, risk, and the other load-bearing values in a decision are inputs to each participant's valuation of the underlying. They shape what a participant thinks a policy is worth and how uncertain they are about it. They are never the unit of price. Earlier drafts claimed that this design choice removes the objection that such values cannot be summed. It does not remove it, and the claim is withdrawn. On the derived path of Section 5 a contribution is a stated transform of an assessed value, so the sum of contributions is a sum of transformed value judgments, and the deeper objection in the blocked-exchange tradition is to the mode of valuation itself, that expressing a judgment about safety as a quantity of a numeraire can change what is being expressed, not to the arithmetic of summation. The design choice narrows the objection to that honest form. Section 13 carries it as a failure mode with a measurement attached, and Appendix A carries it as standing.

5. Participant inputs

For each underlying, each participant i supplies a small set of inputs. The mapping to the inputs of the Black and Scholes (1973) and Merton (1973) valuation is a deliberate analogy, adopted because it makes the sensitivity outputs in Section 9 well defined and legible, and Section 9 states which of those outputs are exact derivatives of the stated aggregation and which are heuristics.

Assessed value, S_i(K). Participant i's assessed value of the policy's outcome for them, at each strike. This is where the load-bearing values enter. A participant who reads a term such as safety differently from the rest of the group will assess a different S.

Reservation strike, K_i*. The strike below or above which participant i prefers their outside option. At strikes on the wrong side of the reservation, participant i contributes nothing. This is the exit parameter, and it is required. Two honest notes attach to it. Declaring it is free and unverifiable, and Section 6 explains why a strategic participant may be tempted to misdeclare it. And a refusal to price is not exit; a participant who declines to translate a protected value into the numeraire records a refusal, a distinct declared input that is counted and reported separately and excluded from the exit statistics, because a protest coded as a preference for an outside option would corrupt both the chain and the validation.

Uncertainty, sigma_i(K) and a_i(K). Participant i's uncertainty about the outcome at each strike, elicited as a range of plausible outcomes at a stated coverage level and a confidence in that range, following the distinction between risk and ambiguity in Ellsberg (1961) and the motivation of the smooth ambiguity model of Klibanoff, Marinacci, and Mukerji (2005). Three limits are stated rather than hidden. The elicited range is not calibrated; interval judgments are documented to be overprecise, and the degree of overprecision depends on the elicitation format (Soll and Klayman, 2004), so cross-participant comparisons of widths are flagged in every output that makes them, and nothing in this mechanism can correct overprecision, because Section 12 forfeits settlement and with it the only general correction device. The confidence input a_i(K) is collected and reported only. It enters no aggregation, no sensitivity, and no decision rule; it is published beside the dispersion surface as the ambiguity surface, an early-warning display and nothing more. And the design does not implement the smooth ambiguity model's machinery, which separates the amount of ambiguity from the attitude toward it; a single confidence number collapses the two, so the mispricing that ambiguity aversion works on the chain is made visible here but not corrected, and it is recorded in Appendix A as standing.

Contribution, c_i(K). The amount of the numeraire participant i commits to contribute if the policy is enacted at strike K. A contribution is a binding pledge, not an expression of enthusiasm, and it is collected only when the policy is enacted. Binding the pledge makes understating costly only to the extent the participant's contribution is pivotal for enactment, a condition Section 6 examines rather than assumes. Participants may supply contributions directly or let the model derive them from S, K*, sigma, the horizons, and the cost of delay. The two paths need not sit on the same behavioral scale, a declared binding number and a model transform of a hypothetical assessment are different objects, so the specification records the elicitation mode per participant and reports the mode mix per strike beside every chain. Contributions are subject to the participant's total budget and to the pricing rule in Section 6.

Blocking contribution, b_i(K). Optionally, the amount participant i would contribute to block the policy at strike K. This prices opposition intensity, and Section 7 states how it enters the decision, a question an earlier draft left unanswered.

Inputs are collected through a stated elicitation protocol rather than free-form entry. Participants compare scenarios, place ranges at the stated coverage level, and confirm the reservation strike by answering whether they would rather have the policy at that setting or their named outside option, with the confirmation procedure's search rule stated, since threshold elicitation carries documented starting-point effects. The protocol is part of the epistemic mapping layer and is codesigned with participants, because the act of translating a value such as fairness into an input is itself an interpretive act and can carry the very latent differences the method is meant to surface. Where a language model assists elicitation it acts before the sealed first round of Section 11, upstream of every anchoring control, so the assisted-versus-unassisted shift in K_i* and S_i(K) is a preregistered measurement in Section 14, not an assurance. Inputs may be revised at any time before the decision horizon. Every revision triggers a recomputation of the chain. The model is deterministic given its logged inputs, so the chain at any moment is fully reproducible from the input log.

6. Aggregation and the pricing rule

The consensus price of the policy at strike K is the aggregate contribution across participants.

P(K) = Σ_i  c_i(K)     for all participants i with K on the acceptable side of K_i*

This aggregate is Lindahl-inspired, and it is not a Lindahl equilibrium. Lindahl (1919) prices a public good through personalized prices, and Foley (1970) proved that a Lindahl equilibrium allocation, personalized prices supporting a common demanded quantity, lies in the core. Nothing here computes personalized prices or clears a common demand, so Foley's theorem does not transfer to this sum, and no core property is claimed for it. That is exactly why Section 10 tests the core directly; if Foley's result transferred, Section 10 would be dead code. The chain P(K), evaluated across all strikes, is the policy's consensus option chain, a name that Section 12 disciplines.

Two rules govern contributions to prevent the aggregate from being captured by intensity alone.

Near-equal budgets. When the numeraire is artificial, each participant receives a budget equal up to the small perturbation that Budish (2011) showed is required, since his approximate competitive equilibrium uses incomes that are unequal but arbitrarily close together precisely because exact equality breaks existence, and its strategyproofness is a large-market property that small groups do not inherit. The artificial numeraire with near-equal budgets is the recommended default. When the numeraire is real, the mechanism changes character, and the specification says so rather than reporting around it. Capacities to pledge differ, so P(K) is censored from below by capacity to pay, the buy rule of Section 7 becomes in part a test of affordability rather than support, N(K) cannot distinguish a participant who declines from one who cannot pay, and a subgroup that cannot afford to fund an alternative cannot register as a blocking coalition in Section 10 however much it objects, so stability verdicts become capacity-bounded. On the real-numeraire path the specification therefore requires the capacity distribution to be recorded before inputs are collected, requires N(K) to be decomposed into declined and unable, reports results raw and budget-normalized, labels every stability verdict capacity-bounded, and refuses the real numeraire where capacity heterogeneity exceeds a threshold the group sets in advance.

Quadratic pricing of intensity. Registering a schedule consumes the participant's budget at a cost that rises with the square of the schedule's magnitude, in the spirit of quadratic voting (Lalley and Weyl, 2018). The charge is levied once per underlying on the strike-density-weighted magnitude of the whole schedule, not per strike, so that refining the strike grid neither raises the charge nor deflates the aggregate; an earlier per-strike reading would have made every number in the specification depend on how finely the designer sampled the dial and would have rewarded concentrating a schedule into a spike, defects reviewers demonstrated at the desk. The welfare claim is stated at its supported strength. Quadratic charging is approximately welfare-optimal in large populations for a binary decision when stated intensity is proportional to true intensity, conditions this setting does not meet, so no welfare theorem is claimed here; the property used is weaker, that marginal influence cost rises linearly with intensity, which dampens capture by any single participant relative to a linear charge. The budget and the pledge are distinct objects. The budget is the influence allowance the quadratic charge draws down; the pledge is the contribution itself, collected in the numeraire on enactment. Keeping the two distinct prevents a unit of the numeraire from counting twice, once as spend and once as voice. The charge's semantics over time are part of the specification, because they decide whether quoting is free. The charge is drawn down when a schedule is registered and is released on revision net of a stated revision cost, so a quote posted and later withdrawn is not free, and Section 11's surveillance report watches the pattern of postings and withdrawals. How the quadratic charge interacts with the shape of the chain is a measured question in the testbed, with the analytical results already in hand recorded in Appendix A.

Lindahl-style aggregation is not strategyproof, and for the goods much of AI governance is about, the standard defense inverts. For an excludable or club-like underlying, a coalition service, a shared tool, a jurisdiction one can leave in Buchanan's (1965) sense, withdrawal removes the benefit, so a binding pledge has teeth and understating genuinely risks losing the policy. For a non-excludable policy enacted over the whole group, a non-contributor still receives the policy, so a binding pledge collected on enactment raises the cost of contributing without conditioning the benefit on it, the best response is to understate except when pivotal, and pivotality vanishes as the group grows. Worse, the exit input becomes an attack surface, since declaring a reservation strike that excludes oneself at the strikes one most wants buys the policy at zero cost while suppressing the chain and inserting a fictitious blocking position into Section 10's search. The specification does not claim a defense it does not have. It states excludability as a scope condition in Section 13, names participation floors with money-back provision and assurance-contract refund devices as the honest instruments for the non-excludable case, as extensions rather than specified mechanisms, and treats the residual gain to misreporting, including reservation-strike misreporting, as a quantity the testbed measures under a stated threat model. Where a sponsor exists, contributions may additionally be matched under the quadratic funding rule of Buterin, Hitzig, and Weyl (2019), under which the matched amount is proportional to the square of the sum of the square roots of contributions; its efficiency is an equilibrium property under that rule's stated assumptions, not a consequence of sincere bidding, its known collusion exposure is bounded by identity verification only to the extent identity verification works, and only the raw sum P(K) faces the buy rule, with the matched amount reported beside it and never substituted silently. One further exposure is stated because the design choice creates it. A quadratic charge makes identity-splitting superlinearly profitable, splitting a contribution across k identities divides its cost by k, so this design is more Sybil-exposed than a linear one, not less, and identity verification is a load-bearing assumption.

Two secondary aggregates are reported alongside P(K), because they answer questions the sum cannot.

Participation count, N(K). The number of participants with a nonzero contribution at strike K, decomposed on the real-numeraire path into declined and unable. This is the direct measure of who would show up, within that stated limit.

Concentration, H(K). The Herfindahl concentration of contributions at K, the sum of squared contribution shares. A chain that clears on the strength of a few large contributors, with N(K) low, is fragile by construction.

7. The buy rule

Let C(K) be the cost, in the numeraire, of enacting the policy at strike K. The group can consider committing at any strike where aggregate willingness meets cost.

Commit at K only if  P(K) ≥ C(K)

This is a total-benefit participation test, and it is not the Samuelson condition, a correction to earlier drafts. Samuelson (1954) states a marginal condition, the sum of marginal rates of substitution equal to the marginal rate of transformation, which selects the efficient level of provision. The analogue of that condition here is the first-order condition at the surplus-maximizing strike, where the slope of P equals the slope of C. The specification reports both objects under their right names, the set of strikes passing the participation test, and within it the strike that maximizes surplus, P(K) minus C(K), whose optimality condition is the Samuelson-inspired one, together with the strike that maximizes participation, N(K). When the surplus-maximizing and participation-maximizing strikes differ, the group has a decision to make, and the difference between them is itself information. A large gap means the surplus-maximizing setting is being carried by a few.

Passing the test at an instant is not sufficient. Section 11 shows that the displayed chain feeds back into inputs, so P(K) can cross C(K) on a transient the mechanism itself created, and real options theory says a group holding a live option to wait does not rationally commit at break-even (McDonald and Siegel, 1986). Commitment therefore requires the condition to hold by a stated margin for a stated dwell time. The margin and the dwell time are declared before inputs are collected, the margin stands in crudely for the option value of waiting rather than pricing it, and the testbed sweeps both rather than treating them as solved.

Where blocking contributions are collected, they enter the decision rather than decorating it, another correction. The specification reports B(K), the aggregate blocking contribution, beside P(K), and any strike where the participation test passes but B(K) is comparable to P(K) is labeled contested and routed to deliberation before commitment. A strike bought over priced opposition is a purchase, not a consensus, and the mechanism says which it is producing.

Across underlyings, chains are compared on quantities that share a unit. Surplus at the selected strike per unit of numeraire spent, participation, concentration, the contested label, and the stability verdict of Section 10. An earlier draft ranked policies by an expected outcome and dispersion read from the Section 8 curve; that ranking subtracted a cost in the numeraire from an expectation in strike units and compared densities over dials that measure different things, it was dimensionally incoherent twice over, and it is withdrawn.

8. The shape of the chain

A priced chain contains more than a set of prices, and this section states honestly what can and cannot be read from it. Earlier drafts claimed that the Breeden and Litzenberger relation recovers, from the chain, the group's implied probability distribution over how the policy will play out. That claim is withdrawn. This section explains why, states the conditional under which a distribution reading would be licensed, and specifies the diagnostics that survive at full strength.

Breeden and Litzenberger (1978) showed that in an option market the second derivative of call prices with respect to strike recovers the prices of state-contingent claims, and the result works because of structure this mechanism does not have. The hypotheses, and their status here, are these.

1. A random variable. The theorem needs a terminal quantity whose realizations share an axis with the strikes, so that a density over strikes is a density over outcomes. Here the strike is a setting the group chooses, not a realization of anything, and no outcome variable is defined unless the group declares one under Section 4.

2. A payoff kernel. The call price is the discounted expectation of a payoff kinked at the strike, and differentiating twice in the strike pulls the density out of that kink. c_i(K) is a pledge conditional on the dial being set to K, not the price of a claim contingent on an outcome exceeding K, so no such representation exists for P.

3. A linear pricing functional. In a market this is no-arbitrage, which also forces call prices to be non-increasing and convex in strike and supplies the boundary conditions that normalize the recovered mass. Section 12 correctly disclaims arbitrage, so nothing forces monotonicity or convexity here, nothing normalizes the curvature, and the chain of a group with an interior consensus is generically single-peaked, rising then falling, which no call chain is.

4. A common measure. Even in a real market the recovered object is a state-price density, probability weighted by marginal utility, not the market's beliefs. Here the weighting is by declared contribution, which mixes belief, preference intensity, budget, and ambiguity attitude, and with no settlement event there is no way to estimate the wedge afterward.

5. A canonical scale. Convexity and curvature are not invariant to monotone relabeling of the dial. One set of declared contributions can be convex on one defensible parameterization of the strike axis and concave on another, so any curvature statement is a statement about the chosen units as much as about the group.

Because hypotheses fail rather than hold, the specification reports the second difference of the raw discrete chain under the name curvature profile, with its sign retained, as a diagnostic of where declared contributions concentrate and where marginal support changes fastest, and it attaches no probability vocabulary to it. Negative curvature around an interior peak is the expected signature of a group with interior agreement, not an inconsistency, and the earlier draft's rule that flagged non-convex regions as internally inconsistent contributions is withdrawn along with its convex-regression repair, which reviewers showed deletes precisely the peak the group most needs to see, returns affine fits whose curvature is identically zero, and manufactures curvature at solver knots that no participant declared. Kinks at reservation crossings are exit boundaries, they are reported as such, and the profile is not differenced across them. The earlier draft's supporting argument that derived schedules are convex because a sum of schedules convex on either side of a kink is convex was a mathematical error, a function convex on either side of a breakpoint is piecewise convex, not convex, and it is retracted.

A distribution reading remains available as a conditional, stated so that a builder can create the conditions deliberately rather than assume them. If the group declares an outcome variable X under Section 4, if contributions are elicited as pledges contingent on the realized X relative to K so that P acquires the kinked-payoff representation the theorem needs, and if the resulting chain is non-increasing and convex in K with declared boundary behavior, then the normalized second difference may be reported under the name commitment-weighted density over X, with the caveat attached that it is willingness-weighted rather than a belief distribution, and with every region where a condition fails excluded. By default none of these conditions holds and nothing of the kind is reported. The abstract of this paper says the same thing, because the honesty must be consistent everywhere.

The diagnostics that carry the section's original purpose, surfacing latent differences before a critical point, are computed from the individual schedules the mechanism already collects, where they are identified, rather than inferred from aggregate curvature, where they are not. Bimodality in an aggregate curve is generated equally by two belief populations, two intensity populations, two clusters of reservation strikes under unanimous beliefs, or two interval-setting habits, so aggregate shape alone cannot say which; the schedules can. For each underlying the specification reports four objects.

1. The exit profile. The empirical distribution of reservation strikes across participants, equivalently the falling edges of N(K). This is a genuine density over the dial, it needs no convexity, no repair, and no grid convention, and its clusters locate the coalitions most likely to leave and where.

2. The ideal-point profile. The empirical distribution of each participant's peak-contribution strike. Its modes are countable, and two modes over a term the group describes as settled is the latent-difference diagnostic of earlier drafts, now stated on an object that identifies it.

3. The camp report. The clustering of full input schedules that Section 10 already runs, reporting the number of camps, their sizes, their separation on the dial, and the strike at which each camp's contribution falls to zero.

4. The dispersion surface. The cross-participant spread of sigma_i(K) at each strike, published with the format caveat of Section 5, and beside it the ambiguity surface of a_i(K), reported and nothing more. The name implied volatility is not used for either, since nothing is inverted through a pricing model; these are declared dispersion and declared confidence.

One reporting rule guards the whole section. Every shape diagnostic is reported only where it exceeds an instrument-effect floor estimated by the two-arm elicitation design of Section 14, in which arms differ in strike grid, presentation order, and elicitation mode; a chain that is flat across the non-exit interior at a rate indistinguishable from the floor is declared uninformative rather than read, since stated willingness to pay is documented to be insensitive to scope (Kahneman and Knetsch, 1992) and the dial is a scope manipulation. Suppressed regions are labeled instrument-limited in the output, because a group whose instrument cannot resolve its own disagreement has learned something too.

9. Sensitivity and coherence outputs

Because each participant's contribution is a stated function of their inputs, the aggregate chain has well-defined sensitivities to those inputs wherever it is differentiable, and the specification reports them under the familiar option names as declared analogies, with three disciplines stated up front. The derivatives here are taken in the strike or in declared inputs, so the delta and gamma below are what an options desk would call dual quantities, not derivatives in an underlying price, of which there is none. The chain is not differentiable in K at reservation crossings, where the summation set itself changes, so delta and gamma are defined between exit boundaries and undefined at them. And the five outputs are not equally rigorous. Delta and gamma are exact derivatives of the stated aggregation where defined; vega is an exact derivative with respect to a declared input; theta and rho are heuristics resting on declared horizons and outside options, are labeled with their calibration status in every report, and carry no asserted shape.

Delta. The slope of P(K) in the strike, between exit boundaries. It tells the group which direction of adjustment in the policy's terms gains the most support.

Gamma. The curvature of P(K) in the strike, which is the same second difference as the curvature profile of Section 8, stated once so the two sections cannot carry incompatible readings of one number. High curvature marks the settings where support changes fastest per unit of term adjustment, which is where redrafting effort concentrates. Earlier drafts additionally called this a fragility or tipping-point measure; that is a dynamical claim about sensitivity of outcomes to perturbation, which curvature in the strike neither implies nor excludes, and it is withdrawn.

Vega. Sensitivity of P(K) to the declared dispersion inputs sigma_i(K), and to those only, since the ambiguity input does not enter the computation. It shows how much aggregate support depends on the declared uncertainty, which is a fact about the inputs; the earlier claim that it prices what a clarification of a term is worth was a value-of-information claim requiring a decision model this mechanism does not have, and it is withdrawn.

Theta. Decay of P(K) with time, as declared by participants through their inputs. The rigorous home of the concept is the real options literature of Section 3, but no valuation model is calibrated here, so the specification asserts only the sign expectation, that declared decay is negative and grows as the validity horizon approaches, and treats the time profile as a quantity the testbed measures. At the validity horizon the chain is void.

Rho. Sensitivity of P(K) to the value of outside options. Outside options enter through the reservation strike, a threshold, so P is a step function of outside-option value, flat between crossings and discontinuous at them. Rho is therefore reported as a scenario sensitivity, the finite change in P under stated shifts in outside-option values, and it is a coarse quantity by construction.

When blocking contributions are collected, the pair P(K) and B(K) invites a market analogy the specification declines. In a market, calls and puts are tied by put-call parity through a forward price and a discount factor, neither of which exists here, so no parity relation is defined and no departure from one can be measured. What is reported instead is the contested-strike diagnostic of Section 7, the strikes where willingness to enact and willingness to block are simultaneously large, which is where the group has not formed a stable view and deliberation belongs before commitment.

10. The stability test

A policy that meets the buy rule can still fail if a subgroup can do better by leaving. The specification applies a core-inspired test to any candidate commitment, and states its logic and its limits with equal care.

For each candidate subgroup, the test asks whether that subgroup's aggregate contribution to some other underlying, at some strike, exceeds what its members receive under the candidate commitment. If such a blocking coalition is found, the commitment is reported as blocked, the coalition is named under the disclosure protocol below, and the alternative it would prefer is shown.

The verdict is one-sided, and the specification never reports the other side. Candidate subgroups come from three generators, the subgroups the epistemic mapping layer identifies, the clustering of input schedules, and a complementarity-driven deviation search seeded by surplus gradients and by strike coverage on alternative underlyings. The third generator exists because clustering groups the similar while a blocking coalition is a complementarity object, a set whose members cover different strikes of an alternative, so a similarity search systematically misses exactly the coalitions that block; earlier drafts specified the searched set inconsistently across sections, and the broadest search governs. Even so, the searched families are a vanishing fraction of the coalition lattice, so a failure to find a blocker is reported as no blocking coalition found within the searched families, with the searched families enumerated, and never as core-stable. Where the mapping layer proposes candidate coalitions, model inference sits upstream of this test, its proposals are logged with provenance, and when the option-set coverage check of Section 13 fails its floor the verdict is reported as not evaluable rather than as stable. On the real-numeraire path every verdict is additionally labeled capacity-bounded, per Section 6.

Core-stable outcomes may not exist at all (Bogomolnaia and Jackson, 2002). When the search finds a blocker for every candidate commitment, the specification reports the smallest uniform relaxation that clears one, defined as the smallest epsilon such that no searched coalition improves on the commitment by more than epsilon per member, and reports which participants bear it. The relaxation standard in the theoretical literature is the approximate core, for which Song and Nguyen (2026) prove a deterministic guarantee in the transfer-free ordinal setting; no comparable guarantee is claimed for this heuristic search, and the epsilon reported is a description of what the search found, not a bound.

Naming a viable breakaway coalition to a room is not a neutral act, it is a coordination device that can manufacture the defection it predicts, and Hirschman's own argument warns that making exit salient can atrophy voice. The disclosure protocol is therefore ratified by the group before inputs are collected, stating who sees a named coalition and when, with the default that facilitators and the coalition's own members see it first and the full room sees it only alongside the deliberation it triggers.

The core test is the difference between an agreement the group endorses and an agreement the group keeps. It is run after the buy rule, not instead of it, because the two answer different questions.

11. Dynamics

Inputs may change at any time, and the chain, the diagnostics, and the stability test are recomputed on each change. The computation is a sum, a set of differences and derivatives, and a coalition search, and all but the last are linear in participants and strikes; the coalition search is bounded by the searched families of Section 10.

The mechanism is itself a complex adaptive system, and this section specifies the loop rather than asserting its management. The state is the participant-by-strike contribution matrix, the displayed chain, and the display buffer. The response map is the participants' revision behavior given the display, parameterized in the testbed by a response gain on the gap between displayed P(K) and C(K). Its sign matters and is stated. Near clearing strikes the free-rider incentive of Section 6 makes the response negative, a participant who sees comfortable surplus withdraws support and one who sees shortfall adds it, and delayed negative feedback is the canonical generator of oscillation, so the observation lag between a revision and its appearance in the display, which an earlier draft offered as a damper, may destabilize instead. The lag is therefore a swept parameter whose oscillation boundary the testbed maps, not an asserted remedy. The update schedule is asynchronous, and the reached state can depend on arrival order, which the testbed probes by perturbed re-runs; with human participants only the single realized trajectory exists, so the input log supports audit of that trajectory and exact recomputation, which is accounting, not counterfactual attribution, since a logged revision made in response to the display is not the cause of the state it responded to.

Four design rules and one honest concession govern disclosure. The first round is collected sealed, so no participant sees an aggregate before committing an initial schedule, at the stated cost that the trajectory's sensitivity concentrates on that one uninformed draw. Individual schedules are never shown; the aggregate chain and participation count are. Revisions appear in the display after the lag. And the concession is that the best deployed evidence on display feedback, the Community Notes result of Section 1, shows that aggregate-only disclosure with latency did not prevent post-display polarization in a system displaying far less than a priced curve, so disclosure level is treated as an experimental arm in Section 14, sealed throughout against aggregate-with-lag against aggregate-without-lag, rather than as a settled choice.

Convergence is monitored, never celebrated. The change in P(K) per revision is reported together with the moving-window variance and lag-one autocorrelation of P(K), the stability of the curvature profile, and a stability-under-reframing check from the two-arm design of Section 14, because a rapidly settling process may be a strongly anchored or cascading one, rising autocorrelation and variance are the standard early warnings of an approaching regime shift, and a diagnostic that cannot distinguish settling from capture is not a diagnostic. The buy rule's dwell-and-margin requirement of Section 7 exists for the same reason, so that the mechanism does not commit on a transient it manufactured.

The input log carries more than replay. It is a surveillance record, and the specification reports per-participant revision magnitudes, withdrawals near the decision horizon, and the correlation between a participant's revisions and subsequent movement of the displayed aggregate, because Section 6's revision cost makes quoting expensive but detection is what makes manipulation risky. The log also carries, as first-class logged inputs with provenance, every generated artifact the mechanism consumes, the candidate underlyings, the strike grid, the parsed load-bearing terms, the proposed outside options, and the mapping layer's candidate coalitions, each with the generating system's version identifier, decoding configuration, seed, prompt hash, candidate count, and filtering rule, together with the contestation channel's record of challenges to any of them. Determinism holds in the stated sense, the chain at any moment is fully reproducible from the log, including those generated inputs.

12. What this method is not

It is not a market. No one trades, so there is no arbitrage, and the prices are declared contributions passed through stated rules rather than prices discovered by exchange. The specification calls them consensus prices and claims no no-arbitrage properties.

It is not a prediction market. Hanson's market scoring rule aggregates beliefs into a price by rewarding traders against a settled outcome, and a scoring-rule market over a continuous outcome natively yields a group probability distribution disciplined by that settlement. Nothing here settles. A contribution is a priced commitment, never a scored forecast, so the mechanism inherits none of a scoring rule's incentive or calibration properties, produces no distribution, and, having forfeited settlement, cannot correct the overprecision of its declared uncertainty inputs.

It is not Black and Scholes pricing. The valuation inputs are mapped to the Black and Scholes inputs because the analogy makes the sensitivities legible, but the chain is not derived by replication and hedging, since there is no underlying asset to hold.

It is not a Lindahl equilibrium. The aggregate is a sum of declared schedules, no personalized prices are computed and no common demand is cleared, so Foley's core theorem does not apply to it, which is why Section 10 tests stability directly.

It is not the Samuelson condition. The buy rule is a total-benefit participation test; the Samuelson-inspired marginal condition appears only at the surplus-maximizing strike, and the two are named separately throughout.

It is not a Breeden and Litzenberger density. The curvature of the chain is a diagnostic of declared contributions, not a probability distribution over outcomes, for the enumerated reasons in Section 8, and no probability vocabulary attaches to it anywhere in this specification unless the conditional of Section 8 is deliberately constructed and declared.

It does not escape the impossibility results, and it names the ones that bind. Arrow's theorem (Arrow, 1951) and the Gibbard and Satterthwaite theorem (Gibbard, 1973; Satterthwaite, 1975) constrain ordinal aggregation without transfers; a cardinal mechanism with a numeraire is not their direct subject, and it earns nothing by the observation, because the public goods impossibility of Green and Laffont (1977) does bind it, no mechanism is simultaneously efficient, dominant-strategy incentive compatible, and budget balanced for public goods. The claim is narrower and survives at that strength. Making exit a first-class input changes which outcomes are selected and filters out some that are unstable under voluntary participation, and nothing here is incentive compatible.

It does not put a language model in the pricing arithmetic, and it does not pretend that settles the question. Aggregation, the buy rule, the curvature profile, the sensitivities, and the core test are deterministic functions of logged inputs and contain no model inference. The domain of those functions, the candidate underlyings, the strike grid, the parsed terms, the proposed outside options, and part of the candidate coalition set, may be model-generated, which is agenda power, so those artifacts are logged with provenance, checked for coverage, and open to challenge through the contestation channel, and the language-model integration of the companion framework is designed and largely untested, a maturity statement that applies wherever models appear in this specification.

It does not replace deliberation, and the honest statement of the relationship is that it supplies an agenda, a decision rule, and a stability verdict, which is more than a neutral diagnostic. The outputs, the commitment range, the participation count, the disagreement report, the contested strikes, and the named coalitions, are things people should know before they decide. The group ratifies or contests the agenda, deliberates where the mechanism flags contest, and decides; a reason field attached to each schedule, disclosed in aggregate, lets something other than a price move the room. Why an output of this procedure should bind those who lose by it is a question of political legitimacy the specification does not answer, and Appendix A records it as standing.

13. Assumptions, scope conditions, and failure modes

The method rests on three assumptions.

Sincerity. Participants can state a reservation strike and an assessed value with some sincerity. Near-equal budgets and quadratic pricing reduce the payoff to misreporting; nothing eliminates it, and Section 6 states where the incentive inverts.

Numeraire. A single numeraire is acceptable to the group as the thing it spends. The artificial equal-budget numeraire is the default; the real-numeraire path is admissible only under the capacity conditions of Section 6.

Outside options. Outside options can be named or generated and ratified. Where they cannot, the reservation strike degenerates to a simple threshold, the core test cannot be run, and the mechanism loses its differentiating input, a limit the abstract's claims are bounded by.

Four scope conditions bound the class of decisions the mechanism should be run on.

1. Excludability. The underlying is club-like, a benefit withdrawal actually removes, or the deployment adds a participation floor with money-back provision or an assurance-contract refund device. For a fully non-excludable policy the pledge defense inverts, per Section 6.

2. Genuine exit. A named, feasible outside option exists for participants. Where none does, the core test is disabled and the buy rule is reported as an affordability test, not a measure of support.

3. Revocability. The decision does not price away a right or protection whose holder cannot be compensated in the numeraire, and it does not bind non-participants; decisions of that class are excluded.

4. Capacity parity. On the real-numeraire path, the capacity distribution is measured first and the heterogeneity threshold of Section 6 is respected, else the artificial numeraire is mandatory.

The known failure modes are these.

1. An empty core, which the relaxation report of Section 10 states rather than hides, with its bearers named.

2. Manipulation of contributions, of reservation strikes, and of the shape of the chain by multi-strike quoting, bounded by the charge semantics and surveilled through the log, measured in the testbed under a stated threat model, and never claimed solved.

3. Unrealistic or missing generated outside options, controlled by provenance, human ratification, the contestation channel, and the coverage audit below, with the core verdict reported as not evaluable when coverage fails its floor.

4. A chain too sparse or too kinked to support shape diagnostics, which the reporting rules of Section 8 refuse to over-read.

5. Feedback-driven oscillation or herding, including the possibility that the observation lag destabilizes rather than damps, monitored by the diagnostics of Section 11 and bounded by the dwell-and-margin rule of Section 7.

6. Elicitation bias, the instrument shaping the inputs, measured by the two-arm instrument-effect floor of Section 14 and by the assisted-versus-unassisted elicitation arm.

7. Motivational crowding and protest response. Offering a price can reduce or distort stated willingness where values are protected, and the price level can carry no information at all (Frey and Oberholzer-Gee, 1997); refusals to price are recorded under their own code per Section 5, and the priced-versus-unpriced comparison of Section 14 measures the frame effect.

8. Scope-insensitive contributions, stated willingness flat across the dial (Kahneman and Knetsch, 1992), caught by the informativeness check of Section 8 and reported as an uninformative chain rather than read.

9. Overprecision of interval inputs, documented and format-dependent (Soll and Klayman, 2004), uncorrectable here for want of settlement, so dispersion outputs carry their cross-participant caveat permanently.

10. Wealth censoring on the real-numeraire path, which Section 6 requires be reported, decomposed, and bounded rather than discovered by the reader.

Option-set coverage is audited, not assumed. The generated outside-option set is compared against a control set elicited without model assistance from a held-out subgroup, reporting recall against the control set and a diversity statistic, with a pre-declared floor. Below the floor, the core verdict is not evaluable and says so.

14. Validation

The method is validated in silico before it is used with people, and this section is a preregistration design, not a result. Two maturity statements bound every claim that follows. The agent population is instantiated from survey data on the agent-based social simulation framework described in the companion white paper, whose validation to Department of Defense verification, validation, and accreditation standards covers its empirical population dynamics under its original intended use; accreditation is scoped to intended use and does not travel to mechanism comparison, no validation of coalition formation or defection dynamics is claimed by either paper, and the framework's language-model integration is designed and largely untested. The comparison described here is therefore a plan whose critical path runs through work that does not yet exist, and every "is measured" of earlier drafts reads "would be measured" here.

The design is factorial, because a bundle compared against unbundled baselines attributes nothing. Aggregation rules form the rows, a Schulze-style forced-choice aggregator, a bridging aggregator, generative social choice, a leveled-commitment arm in which exit is priced as a decommitment penalty with no chain (Sandholm and Lesser, 2001), the chain mechanism itself, a random-strike control, and a status-quo control. The Section 10 stability filter forms the columns, applied and not applied to each rule's output, since the filter is detachable and may carry any durability difference by itself. All arms select from a common outcome space, the baselines' candidate menus generated one candidate per strike from the same grid, with an unequalized variant reported secondarily. The tie-break rule converting the chain arm's commitment range into a single commitment is stated and preregistered. The generation configuration for every model-assisted arm, decoding parameters, candidate counts, prompt templates, and seeds, is frozen and preregistered, shared identically across arms, and ablated over at least two independent configurations, with the stated falsifier that a durability ordering that reverses under ablation is no result.

Grading is independent of the graded. Coalition and defection outcomes are scored by an oracle separate from the mechanism's own search, exhaustive on instances small enough to permit it and randomized with a stated budget otherwise; the mechanism's own enumerator is never used to score, and its miss rate against the oracle is itself a reported diagnostic. The agent model is written down, the latent variables, the map from latents to elicited inputs, and the map from latents to exit and arrival decisions, with at least one variable entering the defection map that enters no mechanism's input, so that the selection rule is demonstrably not a sufficient statistic for the exit rule. The agents' exit destinations are drawn from a set disjoint from the alternative set the pricing arm priced over, with the overlap reported. Defection hazards are anchored to documented human behavior in coalition and public-goods settings as direction-and-shape consistency, never calibration.

The outcome measures are durability, the fraction of agreements surviving once exit is allowed, with time to first defection, blocked-commitment frequency under the independent oracle, and the gain to a manipulating agent under a stated threat model naming the manipulator's information set, their number, and whether they collude. Durability is measured under exit and also under entry, since the one deployed measurement of consensus instability, the Community Notes result, is an entry phenomenon, and a mechanism robust to leavers but brittle to arrivers would pass a test that only lets agents leave. Secondary arms measure what the design's own honesty requires, the instrument-effect floor from two elicitation arms differing in grid, order, and mode; the priced-versus-unpriced frame comparison of Section 13; the assisted-versus-unassisted elicitation shift in K_i* and S_i(K); anchoring and oscillation under the disclosure arms of Section 11 with the lag and response gain swept; and participant burden by arm, since this mechanism demands more numeracy than a ranking and the advantage to the calibrated is a channel the baselines do not share. The preregistration fixes the primary endpoint, durability, the estimator and its clustering by scenario, a minimum detectable effect with the planned scenario and seed counts, the multiplicity correction, numeric thresholds in place of "materially more durable" and "without a large loss", and the provenance of held-out scenarios.

The intended-use statement is part of the design. A positive result would license the claim that, in this agent model, under this generator configuration, chain pricing with a stability filter produces more durable agreements than forced-choice baselines; it would license nothing about people until a human study is run. If the pricing layer fails its preregistered threshold, the testbed, its durability metric, and the disagreement report remain the contribution, and the pricing layer is set aside.

15. Relation to the companion white paper

The companion paper, Getting Everyone on the Same Page, proposes a four-layer architecture for surfacing latent epistemological differences, namely an LLM epistemic mapping layer, a provenance layer, an in silico testbed, and a human deliberation layer, with the mapping layer's language-model machinery designed and largely untested. This specification adds a fifth layer between mapping and deliberation. The mapping layer identifies the candidate policies, their load-bearing terms, and the subgroups likely to hold different readings. This layer prices the candidates, publishes the disagreement report that shows where readings diverge, and tests one-sidedly whether any proposed commitment holds once participants are free to leave. The testbed rehearses the result. The deliberation layer decides. Together they replace a single endorsed statement with a priced, exit-aware, and interpretable picture of what a group is actually prepared to do.

Appendix A. Adversarial review log

Draft 0.1 was reviewed from ten expert perspectives. Each objection is recorded with the resolution adopted in draft 0.2 or the reason it was not adopted. Objections are stated at full strength. Section numbers in the resolutions were updated mechanically to draft 0.3's numbering; no objection or resolution was altered in substance. The second-round log below records where these resolutions were later reopened, so the entries A.1 through A.11 are preserved here as the historical record of the first round.

A.1 Behavioral economics.

Participants cannot report an assessed value and an uncertainty as numbers. Judgments are reference dependent and loss averse (Kahneman and Tversky, 1979), and a declared reservation strike will anchor everything that follows.

Inputs are now elicited through scenario comparison and range placement rather than numeric entry (Section 5). The reservation strike is confirmed by a direct choice between the policy at that setting and the named outside option. The elicitation instrument is codesigned and its sensitivity is reported as a failure mode.

A.2 Heuristics and judgment.

Once an aggregate chain is displayed, participants will anchor on it and herd. The mechanism will converge on whatever it showed first.

The first round is sealed, individual schedules are never shown, an observation lag is introduced, and a convergence diagnostic is reported (Section 11). Anchoring is a measured quantity in the testbed.

A.3 Market pricing.

There is no arbitrage, so these are not prices, and calling the dispersion surface implied volatility is a category error. Breeden and Litzenberger requires call prices convex in strike, and a Lindahl sum of arbitrary schedules is not convex, so the density can go negative.

The paper never claimed no-arbitrage and now says so in three places. Convexity is required, either by construction from the valuation function or by convex regression, and regions where it had to be imposed are reported as findings (Section 8). Output labels use implied dispersion; implied volatility appears only as an analogy.

A.4 Decision under uncertainty.

A single sigma conflates risk with ambiguity. Participants facing Knightian uncertainty about a policy are ambiguity averse, and treating their uncertainty as a variance will misprice the chain.

Uncertainty is now two inputs, a range and a confidence in the range, following Ellsberg (1961) and Klibanoff, Marinacci, and Mukerji (2005). Ambiguity is reported separately from dispersion (Section 5).

A.5 Convexity.

Policy outcomes are often threshold-like in the dial, so the chain will have kinks. Gamma read across a kink is not fragility, it is discontinuity. And the quadratic cost interacts with convexity in ways the draft did not state.

Strikes may be piecewise. Gamma is reported only in locally convex regions (Section 9). Convexity violations are reported rather than smoothed away. The interaction of quadratic pricing with chain shape is left as a stated open question in the testbed plan.

A.6 Complex systems.

The mechanism is a complex adaptive system with feedback from output to input. The draft described dynamics as if the chain were a passive readout. Expect path dependence, oscillation, and emergent coalitions the coalition search did not enumerate.

Section 11 now treats the mechanism as a CAS, specifies update rules, lag, and disclosure limits, and adds a convergence diagnostic. The input log allows replay to expose path dependence. Emergent coalitions are found by clustering input schedules, not only by the mapping layer.

A.7 Mechanism design and social choice.

Lindahl aggregation is not strategyproof. Contributions as expressions of willingness invite free riding on a shared good. Quadratic pricing assumes price taking. Equal artificial budgets are Sybil vulnerable. The core may be empty. And the impossibility theorems still bind.

Contributions are now binding pledges collected on enactment (Section 5), which converts free riding into forgone policy. Quadratic funding matching is available with its collusion caveat stated (Section 6). Identity verification is an assumption. Empty core is reported with the smallest relaxation (Section 10). Incentive compatibility is explicitly not claimed and the residual manipulation gain is a testbed metric.

A.8 Large language models.

Generated outside options and parsed terms will carry homogenization and stereotype bias, and an LLM elicitation assistant will nudge inputs toward the model's priors. Any model in the pricing path makes the outputs untraceable.

Section 12 now states that no model inference sits in the pricing path. Generated options carry provenance and require human ratification before they enter the core test. Elicitation assistance is bounded to clarification, not suggestion.

A.9 Multi-agent systems.

A testbed whose agents exit according to the hypothesis being tested proves nothing. Baselines, preregistration, and separation between the population model and the pricing model are required.

Section 14 now states that agent exit behavior is driven by the grounded population dynamics, not the pricing model; that metrics and thresholds are preregistered; and that held-out scenarios are used for the final comparison.

A.10 Deliberative theory.

Pricing commodifies deliberation. Equal budgets still advantage the articulate. Translating fairness into an input is an epistemological act that can smuggle in exactly the latent differences the method claims to surface.

Accepted in part. The paper does not claim to replace deliberation and says so (Section 12). The elicitation instrument is codesigned and its interpretive role is named (Section 5). The advantage to the articulate is real and is not resolved; it is recorded here as an open limitation to be measured against the forced-choice baselines, which share it.

A.11 Objections not resolved

Three objections stand. Incentive compatibility is not achieved and is not claimed. The advantage to articulate participants persists under any elicitation instrument and is shared by every baseline. And the numeraire, real or artificial, remains a normative choice the group must make before the mechanism can run; the method makes the choice explicit but cannot make it for them.

Appendix A, second round

Draft 0.3 was reviewed a second time, by ten independent expert personas, behavioral economics and heuristics; decision making under uncertainty and ambiguity; market microstructure and options pricing; convexity and the mathematics of pricing; mechanism design and social choice; complex adaptive systems; large language models and their failure modes; multi-agent systems and simulation validation; deliberative democracy and political theory; and a skeptical program lead in deliberative technology. The reviews were conducted independently, without sight of one another, each including live verification of the citations load-bearing for its domain, and a supervising editor adjudicated every objection, resolved inter-reviewer conflicts explicitly, and produced this draft 1.0. Entries A.12 onward merge each distinct objection across reviewers, attribute it by persona, and record its disposition; the first-round entries that reviewers reopened receive explicit new dispositions first. Nothing recorded in the first round has been altered or removed.

Second-round dispositions on A.1 through A.11

A.1, reopened by the behavioral economics and decision-under-uncertainty reviewers. The first-round resolution addressed the format of the value and uncertainty inputs while the load-bearing input, the contribution, remained direct numeric entry under a nonlinear budget, ranges were elicited without a coverage level, and the reference-dependence half of the objection was never engaged. New disposition, accepted in part. Section 5 now states the coverage level, states the confirmation search rule for the reservation strike, flags cross-participant width comparisons, and concedes that overprecision is uncorrectable without settlement; Section 8's instrument-effect floor and Section 14's elicitation arms convert the asserted mitigations into measurements. The deeper identification problem is recorded at A.20 and stands.

A.2, reopened by six reviewers. The sealed round leaves designer-supplied anchors, the grid, the order, the budget framing, as the only anchors present; the lag is not a proven damper; the convergence diagnostic scores anchoring as success; and the elicitation assistant anchors upstream of every control. New disposition, accepted. Section 11 now states the loop's sign, treats the lag as a swept parameter, extends the diagnostics, and makes disclosure an experimental arm; Section 14 randomizes grid, order, and mode between arms and preregisters the assisted-versus-unassisted shift; the testbed-tense overreach is corrected everywhere.

A.3, reopened by five reviewers and partially confirmed by two. The label half, implied dispersion as the output name, held and is kept. The substance half failed, requiring convexity was not a response to the finding that the object is not a density, and the convex repair made the situation worse. New disposition, accepted in full at A.12 and A.13; Section 8 is rewritten, the density claim is withdrawn, and the repair is removed.

A.4, reopened by three reviewers against one confirmation. The resolution was a relabeling, the ambiguity input entered nothing. New disposition, accepted. Section 5 now labels a_i(K) collected and reported only, the ambiguity surface is a published output, vega is defined against dispersion only, and the unpriced ambiguity mispricing moves to the standing list.

A.5, reopened by six reviewers. The kink treatment was partly right, but the convexity-by-construction lemma added in draft 0.3 was false, the guard consumed the reportable region, gamma and the density were one number under two readings, and the quadratic interaction was not open but analytically adverse. New disposition, accepted. The lemma is retracted, the repair removed, gamma and the curvature profile unified with one reading, the charge restated grid-invariantly at the schedule level, and the remaining interaction measured in the testbed with the desk results recorded at A.16.

A.6, reopened in part by four reviewers, confirmed in part by two. The CAS framing held; the lag's sign, the ensemble nature of path dependence, the diagnostic's blindness, and the coalition-search contradiction did not. New disposition, accepted. Section 11 is rewritten as a specified loop with extended diagnostics and honest claims about replay; Section 10 resolves the search contradiction in favor of the broadest search.

A.7, reopened on all limbs by four reviewers. The free-riding device inverts for non-excludable goods, the price-taking limb was never engaged, the Sybil resolution restated the problem while the quadratic charge worsens it superlinearly, and the core and impossibility limbs leaned on theorems that do not attach. New disposition, accepted. Sections 6, 12, and 13 state the inversion, the conditions of the quadratic result, the Sybil exposure, the correct impossibility, and the scope conditions; the misattached theorems are corrected at A.14 and A.15.

A.8, reopened in part by four reviewers, confirmed on the pricing-path statement. Ratification filters implausibility, not skew, and is silent on the un-generated option; the clarification-versus-suggestion boundary is one the paper itself dissolves; maturity statements were missing outside Section 14. New disposition, accepted in part. Sections 11, 12, and 13 add provenance logging, the coverage audit with a floor that gates the core verdict, the contestation channel, and maturity statements wherever models appear; the homogenization risk remains a measured question and stands until measured.

A.9, reopened by four reviewers. Module separation is not predicate separation, the framework's accreditation does not cover this use, the baselines were not comparable, and preregistration had no content. New disposition, accepted. Section 14 is rewritten as a factorial preregistration design with an independent grading oracle, a written agent model, disjoint option sets, generator controls, and stated statistical content.

A.10, reopened in part by the deliberative democracy reviewer, confirmed in part by others. The commodification answer in Section 4 was self-defeating on the derived path, the real-numeraire equality gap was unseen, and the shared-by-baselines defense fails for numeracy. New disposition, accepted in part. Section 4 withdraws the overclaim, Section 6 states the censoring, Section 13 adds crowding and scope-insensitivity failure modes with measurements, Section 14 measures participant burden, and the blocked-exchange and legitimacy objections move to the standing list.

A.11, extended. All three first-round standing objections remain standing, restated with more force where the second round sharpened them, and the updated list at the end of this appendix supersedes the count of three.

A.12 The Breeden and Litzenberger reading

Raised, with numerical and analytical demonstrations, by the market microstructure, convexity mathematics, decision under uncertainty, behavioral economics, mechanism design, multi-agent validation, and program lead personas. The relation requires a random variable sharing the strike axis, a kinked payoff kernel, a linear pricing functional, boundary conditions for normalization, and monotone decreasing convex prices; the mechanism has none of these, the strike is a chosen setting rather than a realization, the recovered object even in a market is a state-price density rather than beliefs, the weighting here is by contribution and cannot be divided out without settlement, and curvature is not invariant to relabeling the dial. Accepted in full. The implied-distribution claim is withdrawn from the abstract, Section 3, and Section 8; the second difference is reported as a curvature profile with no probability vocabulary; the distribution reading survives only as the explicit conditional of Section 8 with its failure conditions enumerated; and Section 12 adds "not a Breeden and Litzenberger density".

A.13 The convexity machinery

Raised by the convexity mathematics and market microstructure personas with worked constructions. Derived schedules sum to U-shaped chains whose surplus maximum is forced to an endpoint; realistic direct chains are single-peaked, so most interior strikes carry negative second differences that draft 0.3 branded internally inconsistent; convex regression deletes the peak, returns affine fits with identically zero curvature, or manufactures atoms at solver knots; the exit-boundary guard excludes most or all of the grid at realistic sizes; and the draft 0.3 lemma that a sum of piecewise convex schedules is convex is false. Accepted in full. The convexity requirement, the repair, and the lemma are removed; raw signed curvature is reported; kinks are reported as exit boundaries; non-convexity is no longer called inconsistency.

A.14 Lindahl, Foley, and Samuelson

Raised by the mechanism design, convexity mathematics, market microstructure, and program lead personas. The sum of pledges is not a Lindahl equilibrium, so Foley's core theorem does not transfer, and its invocation contradicted the existence of Sections 10 and 13; the buy rule is a total-benefit test, not Samuelson's marginal condition, whose true analogue the draft wrote one sentence later without noticing. Accepted in full. Sections 6 and 7 are corrected, Section 12 adds the two "is not" items, and the core test carries the stability claim alone.

A.15 Budish, Lalley and Weyl, and Buterin, Hitzig, and Weyl

Raised by the mechanism design, behavioral economics, deliberative democracy, convexity mathematics, and program lead personas. Budish's budgets are unequal but arbitrarily close, with strategyproofness only in the large; the quadratic voting result is approximate, large-population, binary, and conditional on proportional intensity reporting; quadratic funding's matched amount is proportional to the square of the sum of square roots and its efficiency is an equilibrium property, not a consequence of sincerity; and the draft never said which aggregate faces the buy rule. Accepted in full. Section 6 states all four correctly, adopts near-equal budgets, drops the welfare-optimality claim to its supported strength, and fixes the raw sum as the only aggregate the buy rule sees.

A.16 The quadratic charge and the grid

Raised by the mechanism design, market microstructure, and convexity mathematics personas, who showed that a per-strike quadratic charge makes the aggregate fall roughly as the inverse square root of the grid size, makes concentration the dominant declaration strategy, and penalizes flexible participants against inflexible ones. Accepted in part. The charge is restated at the schedule level with strike-density weighting, which removes the grid dependence and part of the concentration incentive; the residual interaction of the charge with chain shape, including its penalty on breadth, is preregistered as a testbed measurement with these desk results recorded as the expectation to test.

A.17 Multi-strike quoting and shape manipulation

Raised by the market microstructure persona. Only one strike settles, so quotes elsewhere are free unless the charge's revision semantics say otherwise, which draft 0.3 did not; anonymous aggregate display is the spoofing-friendly condition; and a butterfly-shaped misreport moves the reported curvature by multiples while barely moving the level. Accepted in part. Section 6 states the charge semantics with a revision cost, Section 11 adds the surveillance report over the input log, and Section 14 adds a red-team arm under a stated threat model. The manipulability of shape as distinct from level is recorded on the standing list.

A.18 Free riding, excludability, and the exit primitive

Raised by the mechanism design and program lead personas. For non-excludable policy goods a binding pledge collected on enactment does not price the benefit a non-contributor still receives, so the draft's primary defense inverts, and the reservation strike becomes the cheapest input to misreport. Accepted. Section 6 states the inversion and the misreporting attack, the "primary defense" language is withdrawn, excludability becomes scope condition one of Section 13 with participation-floor and assurance-contract devices named as the honest instruments for the non-excludable case, and the Nash-implementation literature for Lindahl allocations is acknowledged as the place a builder would go, out of scope for this thought experiment. The residual limitation, that much of AI governance is non-excludable, stands.

A.19 Wealth censoring on the real-numeraire path

Raised by the deliberative democracy persona as a formal property, the buy rule and the core test become capacity-to-pay weighted, N(K) cannot distinguish declining from unable, and a subgroup too poor to fund an alternative cannot register a block, so stability verdicts self-certify against the least resourced. Accepted. Section 6 states the censoring, makes the equal-budget artificial numeraire the recommended default, requires capacity recording, N(K) decomposition, capacity-bounded labels on stability verdicts, and refusal of the real numeraire above a pre-set heterogeneity threshold; Section 13 makes capacity parity a scope condition.

A.20 Behavioral identification of the inputs

Raised by the behavioral economics and decision-under-uncertainty personas. P(K) is a preference-and-budget-weighted aggregate under reference dependence and loss aversion, stated willingness to pay is scope insensitive for exactly this class of good, protest refusals were coded as exit, the direct and derived contribution paths sit on different calibration scales, and interval elicitation is overprecise in a format-dependent way that settlement-free mechanisms cannot correct. Accepted in part. Section 5 adds the refusal code, the coverage level, and the mode-mix report; Section 8 adds the informativeness check and the instrument-effect floor; Section 13 adds crowding, scope insensitivity, and overprecision as failure modes with citations; Section 14 adds the two-arm instrument design and the priced-versus-unpriced comparison. What cannot be repaired is recorded on the standing list, no elicitation fix inside this mechanism can identify a belief object without settlement.

A.21 The ambiguity input

Raised by the decision-under-uncertainty, convexity mathematics, and program lead personas. a_i(K) entered no aggregation, sensitivity, or decision rule, so the A.4 resolution was a relabeling; and a confidence scalar collapses the ambiguity and ambiguity attitude that the cited smooth ambiguity model exists to separate. Accepted. The input is now explicitly collected and reported only, published as the ambiguity surface, the smooth-ambiguity machinery is explicitly not implemented, and the unpriced ambiguity mispricing, which systematically affects long-horizon poorly characterized policies, stands unresolved.

A.22 The empirical case in Section 1

Raised convergently by the program lead, multi-agent validation, complex systems, deliberative democracy, LLM failure modes, and mechanism design personas, with live verification. The Holliday, Kristoffersen, and Pacuit citation was deployed backwards, Schulze was not tested, Instant Runoff was found resistant under limited information, and Condorcet methods were the least manipulable of the eight tested. The Chuai, Lenzini, and Pröllochs figure measures polarized re-rating by arriving or staying contributors, entry and voice rather than exit, and the paper is now published in the Proceedings of the ACM Web Conference 2026. The minority-overweighting finding belongs to Tessler et al. 2024, whose authors read it favorably, and the "artifact of the rule" gloss was the draft's own. The Bogomolnaia and Jackson "routinely blocked" frequency claim was unsupported by a theory paper whose message is that the core may be empty. Accepted in full. Section 1 is rewritten, each source is characterized at its verified strength, the Chuai result is reinterpreted as evidence that bridged consensus is unstable under display and entry, its display-feedback lesson is absorbed into Section 11, and the section now concedes that no field measurement of exit from a deployed deliberation system exists.

A.23 Feedback dynamics

Raised by the complex systems persona and reinforced by the decision-under-uncertainty and microstructure personas. The displayed-surplus loop is negative under the paper's own free-riding incentive, delayed negative feedback is the canonical oscillator, so the lag may destabilize; the buy rule could fire on a transient; the convergence diagnostic maps settling, cascading, and critical slowing down to the same reading; and equilibrium multiplicity under two-sided exit means the dynamics do the selecting. Accepted. Section 11 specifies the state, the response map and its sign, and the extended diagnostics; the lag is a swept parameter; disclosure is an experimental arm; Section 7 adds the dwell-and-margin rule; fixed-point existence is recorded on the standing list.

A.24 The coalition search

Raised by the complex systems, multi-agent validation, program lead, and LLM failure modes personas. Clustering is a similarity operator applied to a complementarity object; the core verdict is one-sided and was reported as two-sided; three passages specified the searched set inconsistently; the empty-core relaxation was attributed to Fain, Goel, and Munagala, whose Lindahl equilibrium is always a core solution in their setting, while the approximate-core guarantee in the paper's own reference list belongs to Song and Nguyen; and model inference sits inside the candidate set while Section 12 said the core test contained none. Accepted. Section 10 adds the complementarity-driven generator, states one-sidedness and search coverage, states the relaxation metric as a description rather than a bound with the attribution corrected, gates the verdict on option-set coverage, and Section 12's claim is restated as a claim about the arithmetic with the upstream model inference named.

A.25 The validation design

Raised by the multi-agent validation persona and reinforced by the LLM failure modes and complex systems personas. The comparison confounded five differences at once, graded the treatment arm with the coalition oracle it filters against, had no common outcome space, no tie-break rule, preregistration in name only, an accreditation claim scoped to a different intended use, present-tense claims about an unbuilt testbed, a generator confound across three of four arms, and an option-set leakage channel; and the durability evidence gap, that the one deployed measurement concerns entry rather than exit, was unaddressed. Accepted in full. Section 14 is rewritten as a factorial preregistration design with an independent grading oracle, disjoint exit-option sets, frozen and ablated generator configurations, a written agent model with a non-elicited defection variable, stated statistical content, durability measured under exit and entry, an intended-use statement, and uniform conditional tense.

A.26 Language-model agenda control

Raised by the LLM failure modes and deliberative democracy personas. The model-free pricing arithmetic operates on a model-authored domain, the underlyings, strike grid, parsed terms, outside options, and candidate coalitions; generated-set coverage was unmeasured with an omission bias pointing at "core stable"; the elicitation assistant anchors upstream of the sealed round; and no participant could contest the axis. Accepted in part. Sections 11 through 13 make generated artifacts logged inputs with provenance, add the coverage audit with a floor gating the core verdict, add the contestation channel and the strike axis as a contestable input, preregister the elicitation-nudge measurement, and add maturity statements. Homogenization of generated option sets remains a measured question and stands until measured.

A.27 Commodification, deliberation, and legitimacy

Raised by the deliberative democracy persona. The Section 4 answer to commodification was self-defeating on the derived path and misidentified the objection; motivational crowding evidence shows the price frame can halve stated willingness and carry no dose response; the anti-herding devices suppress the mutual influence that constitutes deliberation; naming blocking coalitions is a coordination device for defection; Hirschman's own argument warns that salient exit atrophies voice; and the paper never states why its output binds losers. Accepted in part. Section 4 withdraws the overclaim; Section 13 adds the crowding failure mode with the priced-versus-unpriced arm; Section 12 restates honestly what the mechanism supplies, adds the reason field and the contestation channel; Section 10 adds the ratified disclosure protocol and acknowledges the Hirschman tension. The blocked-exchange objection and the legitimacy question stand unresolved.

A.28 Novelty overstatements

Raised by the market microstructure, mechanism design, and multi-agent validation personas. Scoring-rule markets natively produce group-implied distributions, with settlement; leveled commitment contracts already price exit in multi-agent systems; and "none prices the decision as a curve" overreached against mechanisms that elicit demand across levels. Accepted. Section 3 concedes the scoring-rule distribution plainly, cites Sandholm and Lesser, and narrows the claim to the specified combination, a chain of binding commitments as the deliverable, schedule-level diagnostics, and the exit-durability standard.

A.29 The sensitivity outputs

Raised by the market microstructure, convexity mathematics, and program lead personas. Delta and gamma are dual quantities; gamma was the Section 8 density under a second, incompatible reading; theta's asserted shape had no support; rho differentiates a step function; vega's target input was ambiguous; and put-call parity was invoked with no forward and no discount factor. Accepted. Section 9 is rewritten, one reading for the curvature, sign-only theta, scenario rho, dispersion-only vega, and the parity diagnostic replaced by the contested-strike report.

A.30 References and guardrails

Raised across personas and audited exhaustively by the program lead. Two reference entries shipped bracketed placeholders; preprints lacked version numbers and revision dates; Arrow, Gibbard, Satterthwaite, and Schulze were named with no entries; Chuai was still flagged preprint after publication; the about-the-author section deviated from the sanctioned background; Section 15 contained a colon-constructed sentence; and maturity statements were missing where models appear. Accepted in full. The reference list is rebuilt with every entry verified against a live record, Mitrović and Song and Nguyen resolved, versions and dates recorded for all preprints, entries added for Arrow, Gibbard, Satterthwaite, Schulze, Green and Laffont, Frey and Oberholzer-Gee, Kahneman and Knetsch, Soll and Klayman, and Sandholm and Lesser, the published Chuai record cited, the author section corrected, and the colon sentence recast.

A.31 Objections standing unresolved

Ten objections stand after two rounds.

1. Incentive compatibility is not achieved and is not claimed, and within the stated scope no incentive-compatible repair is specified; assurance-contract and participation-floor devices are named, not designed.

2. The advantage to articulate participants persists under any elicitation instrument, and this mechanism additionally advantages the numerate and well calibrated, a channel the forced-choice baselines do not share; Section 14 measures it, nothing resolves it.

3. The numeraire remains a normative choice the group must make before the mechanism can run, and the choice decides whether the mechanism polls preferences or means; the equal-budget artificial default narrows the exposure without settling it.

4. No belief object is identified. Without a settlement event the curvature of the chain cannot be read as belief, overprecision in the declared uncertainty cannot be corrected, and no elicitation repair inside this mechanism changes either fact.

5. Ambiguity is collected and reported but never priced, so ambiguity-averse mispricing, largest for long-horizon and poorly characterized policies, stands.

6. The shape of the chain is manipulable by coordinated multi-strike quoting even where the level is honest; the surveillance report detects patterns, it does not prevent them.

7. For non-excludable underlyings the exit primitive loses incentive content, and much of AI governance is non-excludable; the scope conditions exclude those cases rather than solving them.

8. Motivational crowding, that pricing a civic commitment can change the commitment, is measured by a preregistered comparison and is not resolved by any design device here.

9. The legitimacy question, why an output of this procedure should bind those who lose by it, is not answered.

10. Existence of a fixed point of the revision dynamics is not established, and the core verdict is one-sided by construction, so stability can be evidenced but never certified.

References

Arrow, K. J. (1951). Social choice and individual values. Wiley.

Bakker, M. A., Chadwick, M. J., Sheahan, H. R., Tessler, M. H., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M., and Summerfield, C. (2022). Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems, 35, 38176 to 38189.

Binmore, K., Shaked, A., and Sutton, J. (1989). An outside option experiment. Quarterly Journal of Economics, 104(4), 753 to 770. https://doi.org/10.2307/2937866

Black, F., and Scholes, M. (1973). The pricing of options and corporate liabilities. Journal of Political Economy, 81(3), 637 to 654. https://doi.org/10.1086/260062

Bogomolnaia, A., and Jackson, M. O. (2002). The stability of hedonic coalition structures. Games and Economic Behavior, 38(2), 201 to 230. https://doi.org/10.1006/game.2001.0877

Breeden, D. T., and Litzenberger, R. H. (1978). Prices of state-contingent claims implicit in option prices. Journal of Business, 51(4), 621 to 651. https://doi.org/10.1086/296025

Buchanan, J. M. (1965). An economic theory of clubs. Economica, 32(125), 1 to 14. https://doi.org/10.2307/2552442

Budish, E. (2011). The combinatorial assignment problem: Approximate competitive equilibrium from equal incomes. Journal of Political Economy, 119(6), 1061 to 1103. https://doi.org/10.1086/664613

Buterin, V., Hitzig, Z., and Weyl, E. G. (2019). A flexible design for funding public goods. Management Science, 65(11), 5171 to 5187. https://doi.org/10.1287/mnsc.2019.3337

Chuai, Y., Lenzini, G., and Pröllochs, N. (2026). Consensus stability of Community Notes on X. Proceedings of the ACM Web Conference 2026, 8885 to 8896. https://doi.org/10.1145/3774904.3792987. Also arXiv:2601.14002 (v1, 20 January 2026).

Dixit, A. K., and Pindyck, R. S. (1994). Investment under uncertainty. Princeton University Press.

Ellsberg, D. (1961). Risk, ambiguity, and the Savage axioms. Quarterly Journal of Economics, 75(4), 643 to 669. https://doi.org/10.2307/1884324

Fain, B., Goel, A., and Munagala, K. (2016). The core of the participatory budgeting problem. In Web and Internet Economics (WINE 2016), Lecture Notes in Computer Science 10123, 384 to 399. Springer. https://doi.org/10.1007/978-3-662-54110-4_27

Fish, S., Gölz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and Wüthrich, M. (2024). Generative social choice. Proceedings of the 25th ACM Conference on Economics and Computation, 985. https://doi.org/10.1145/3670865.3673547. Full version at arXiv:2309.01291 (preprint, v3, revised 5 March 2025).

Foley, D. K. (1970). Lindahl's solution and the core of an economy with public goods. Econometrica, 38(1), 66 to 72. https://doi.org/10.2307/1909241

Frey, B. S., and Oberholzer-Gee, F. (1997). The cost of price incentives: An empirical analysis of motivation crowding-out. American Economic Review, 87(4), 746 to 755.

Garg, N., Goel, A., and Plaut, B. (2021). Markets for public decision-making. Social Choice and Welfare, 56(4), 755 to 801. https://doi.org/10.1007/s00355-020-01298-4

Gibbard, A. (1973). Manipulation of voting schemes: A general result. Econometrica, 41(4), 587 to 601. https://doi.org/10.2307/1914083

Green, J., and Laffont, J. J. (1977). Characterization of satisfactory mechanisms for the revelation of preferences for public goods. Econometrica, 45(2), 427 to 438. https://doi.org/10.2307/1911219

Hanson, R. (2003). Combinatorial information market design. Information Systems Frontiers, 5(1), 107 to 119. https://doi.org/10.1023/A:1022058209073

Hanson, R. (2013). Shall we vote on values, but bet on beliefs? Journal of Political Philosophy, 21(2), 151 to 178. https://doi.org/10.1111/jopp.12008

Hirschman, A. O. (1970). Exit, voice, and loyalty: Responses to decline in firms, organizations, and states. Harvard University Press.

Holliday, W. H., Kristoffersen, A., and Pacuit, E. (2025). Learning to manipulate under limited information. Proceedings of the AAAI Conference on Artificial Intelligence, 39(13), 13915 to 13925. https://doi.org/10.1609/aaai.v39i13.33522

Kahneman, D., and Knetsch, J. L. (1992). Valuing public goods: The purchase of moral satisfaction. Journal of Environmental Economics and Management, 22(1), 57 to 70. https://doi.org/10.1016/0095-0696(92)90019-S

Kahneman, D., and Tversky, A. (1979). Prospect theory: An analysis of decision under risk. Econometrica, 47(2), 263 to 291. https://doi.org/10.2307/1914185

Klibanoff, P., Marinacci, M., and Mukerji, S. (2005). A smooth model of decision making under ambiguity. Econometrica, 73(6), 1849 to 1892. https://doi.org/10.1111/j.1468-0262.2005.00640.x

Lalley, S. P., and Weyl, E. G. (2018). Quadratic voting: How mechanism design can radicalize democracy. AEA Papers and Proceedings, 108, 33 to 37. https://doi.org/10.1257/pandp.20181002

Lindahl, E. (1919). Die Gerechtigkeit der Besteuerung. Gleerup, Lund. Partial translation as Just taxation, a positive solution, in R. A. Musgrave and A. T. Peacock (Eds.), Classics in the Theory of Public Finance (1958), 168 to 176. Macmillan.

McDonald, R., and Siegel, D. (1986). The value of waiting to invest. Quarterly Journal of Economics, 101(4), 707 to 727. https://doi.org/10.2307/1884175

Merton, R. C. (1973). Theory of rational option pricing. Bell Journal of Economics and Management Science, 4(1), 141 to 183. https://doi.org/10.2307/3003143

Mitrović, D. (2024). Pre-electoral coalition agreement from the Black-Scholes point of view. Scientific Reports, 14, article 3227. https://doi.org/10.1038/s41598-024-53674-0

Osborne, M. J., and Rubinstein, A. (1990). Bargaining and markets. Academic Press.

Ponsatí, C., and Sákovics, J. (1998). Rubinstein bargaining with two-sided outside options. Economic Theory, 11(3), 667 to 672. https://doi.org/10.1007/s001990050208

Samuelson, P. A. (1954). The pure theory of public expenditure. Review of Economics and Statistics, 36(4), 387 to 389. https://doi.org/10.2307/1925895

Sandholm, T. W., and Lesser, V. R. (2001). Leveled commitment contracts and strategic breach. Games and Economic Behavior, 35, 212 to 270. https://doi.org/10.1006/game.2000.0831

Satterthwaite, M. A. (1975). Strategy-proofness and Arrow's conditions: Existence and correspondence theorems for voting procedures and social welfare functions. Journal of Economic Theory, 10(2), 187 to 217. https://doi.org/10.1016/0022-0531(75)90050-2

Schulze, M. (2011). A new monotonic, clone-independent, reversal symmetric, and condorcet-consistent single-winner election method. Social Choice and Welfare, 36(2), 267 to 303. https://doi.org/10.1007/s00355-010-0475-4

Soll, J. B., and Klayman, J. (2004). Overconfidence in interval estimates. Journal of Experimental Psychology, Learning, Memory, and Cognition, 30(2), 299 to 314. https://doi.org/10.1037/0278-7393.30.2.299

Song, H., and Nguyen, T. (2026). Ordinal Lindahl equilibrium for voting. arXiv:2603.04312 (preprint, v1, 4 March 2026).

Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., Evans, G., Campbell-Gillingham, L., Collins, T., Parkes, D. C., Botvinick, M., and Summerfield, C. (2024). AI can help humans find common ground in democratic deliberation. Science, 386(6719), eadq2852. https://doi.org/10.1126/science.adq2852

Tiebout, C. M. (1956). A pure theory of local expenditures. Journal of Political Economy, 64(5), 416 to 424. https://doi.org/10.1086/257839

About the author

Stephen Lieberman is the founder of Paramerge, an AI safety and governance practice, and home of the Real-World AI Governance Center, the PAN Lab, and the Oversight simulation. At the Naval Postgraduate School he founded and led the Complex Adaptive Social System (CASS) framework, an agent-based social simulation validated to Department of Defense VV&A standards. His work on this specification draws on more than a decade of research applying complex adaptive systems methods to live derivatives markets. He can be reached at stephen@paramerge.com.

The companion paper, Getting Everyone on the Same Page, argues for surfacing the latent differences in meaning that this mechanism asks a group to price.

All Paramerge white papers

Contact Paramerge