Notes on:

Beliefs Over Contracts

Francis Annan & Collin B. Raymond
Working paper
27 July 2026
development · contracts · firms
Talk · Paper · Transcript
Written by Fable 5

Part of NBER Summer Institute 2026 — Development Economics

Francis Annan (UC Berkeley) and Collin Raymond (Cornell), presented by Annan at NBER Summer Institute Development Economics, July 27, 2026. Paper: March 2026 revision from the author’s site. Timestamps refer to the session video.

The paper begins with a dinner. Four years ago Annan found himself eating with the CEO of MTN Mobile Money — the firm whose agents handle deposits and withdrawals for most of Ghana (90% market share) — and asked why millions of agents across the country were all paid on the same wage structure. The CEO’s answer is the whole paper in miniature: “Great question. I keep making money, so why do I want to change it?” Contract theory assumes the principal knows the answer — models let the agent be biased, never the principal. This paper asks, with a national-scale experiment and a matched beliefs elicitation, whether the people who choose contracts can actually predict which contracts work. The title gives away the punchline: managers hold beliefs over contracts, and the beliefs are only loosely attached to reality.

The machinery is serious. A baseline survey of ~6,000 agents in ~700 communities nationwide; ~500 managers up and down MTN’s hierarchy; full administrative transaction records for every agent point; a nationwide audit study (mystery shoppers checking whether agents responded to new incentives by illegally marking up fees — they didn’t, much); and an endline. MTN built an in-house research team around the project — they were told, credibly, that the results would change compensation policy. The experiment randomized 425 market-communities across nine regions into five contracts for three months (twelve weekly pay cycles), each engineered to be expenditure-equivalent ex ante — same expected cost to the firm, calibrated on each agent’s own transaction history — so the arms differ only in incentive shape and risk: the status-quo simple linear piece rate (~4% commission, the industry standard); a threshold contract (large bonus ≈1.5× the piece rate for hitting your own historical average in a week; miss it and “you go home without any pay”); franchising (pay an upfront platform fee, keep boosted commissions — the Amazon model, and where the industry says it’s heading); a tournament (weekly rank-order pay within homogeneous local groups, total pot fixed); and a flat wage (paid regardless of output — the arm MTN kept asking to remove mid-experiment, and Annan kept saying “just wait”). Agents could quit any arm at any time, reverting to the status quo — crucially, they stay observable in the admin data, so attrition is an outcome, not a hole.

Before randomization, managers ranked the five contracts — after hours of patient explanation, “almost a five-hour type exercise” — on predicted revenue and predicted dropout risk, incentive-compatibly (higher-ranked contracts got later randomization slots, tying beliefs to consequences). The predictions: threshold best for revenue, flat wage worst, linear safest on dropout. And in a lovely aside, when asked which contract they’d choose, about 40% of managers picked something other than what they believed would maximize revenue — whatever the principal is maximizing, it isn’t always the model’s objective function. Asked to explain their rankings, managers talked about incentives and — tellingly — simplicity. Risk, the concept our contract theory revolves around, “is barely mentioned.”

Then the world voted:

Total revenue is highest under tournaments
Slide at 06:11:27: tournaments raise total revenue about 20% relative to the simple linear benchmark; the flat wage is worst; the predicted winner — threshold — underperforms the status quo.

The best-performing contract was the tournament — up roughly 20% on revenue versus the status quo, with similar participation, gains roughly shared between firm and agents — a contract managers ranked middling. Their predicted champion, the threshold, generated the highest dropout (concentrated in the first weeks: agents who hate a contract leave fast) and mediocre revenue. They did correctly identify the losers: the flat wage was indeed worst, with the damage driven entirely by owner-operated shops — an owner paid regardless of output simply stops showing up, while a hired worker still does, which is about as clean a micro-illustration of incentive theory as the data could offer. Franchising, Annan’s own prior favorite, landed statistically indistinguishable from the status quo (theory’s revenge: franchising loads all the risk onto the agent, and these agents face plenty). Overall, the correlation between managerial rankings and realized performance is about 0.2. Managers are good at spotting bad contracts and bad at identifying the best one — an asymmetry Annan attributes partly to the simplicity heuristic: complex contracts (threshold, tournament) carry the largest prediction errors.

The effort data add a genuinely novel wrinkle. A labor-supply index (days worked, hours, early opening, late closing) rises under threshold and tournament — and here the admin data, which keep flowing after the experiment ended, show something the survey couldn’t: after incentives were removed, every contract’s labor supply converged back to the linear baseline except the tournament’s, which stayed elevated. Three months of ranked competition seems to have installed a work habit that outlived the prize money — “micro-foundations,” Annan suggests, “for persistence in labor supply,” or possibly just evidence that people keep caring who’s first even when it stops paying.

Labor supply: habit formation, not intertemporal reallocation
Slide at 06:16:12: admin-data labor supply rises under tournament during the experiment and stays elevated after incentives end, while other contracts converge back.

One more finding deserves its name: the ivory-tower trade-off. Senior HQ executives beat junior field managers at predicting performance (the incentive-compatibility side — they have the data, they “look at markets from the sky”) but junior managers beat them at predicting dropout (the participation side — they know which agent is one bad week from quitting). Sophistication about the IC constraint and sophistication about the IR constraint live in different offices.

The concluding heresy is aimed at the foundations. Behavioral contract theory has spent two decades on sophisticated firms exploiting biased consumers and workers. Here the principal is the unsophisticated party — a large, profitable, data-rich firm whose managers systematically misrank their own incentive schemes, leaving 20% of revenue on the table because the status quo was simple and nobody had run the experiment. The CEO’s dinner-table logic — I keep making money, why change? — turns out to be not an anomaly but the equilibrium: a contract does not have to be optimal to persist. It only has to be good enough that nobody ever finds out.