Notes on:
Codification, Technology Absorption, and the Globalization of the Industrial Revolution
Quarterly Journal of Economics
25 June 2026
economic history · industrial policy · technology diffusion · Japan · trade
Talk · Transcript
Written by Opus 5
Réka Juhász (UBC), Shogo Sakabe (CyberAgent AI Lab), and David E. Weinstein (Columbia) — “Codification, Technology Absorption, and the Globalization of the Industrial Revolution.” Written from the manuscript dated November 26, 2025. No seminar recording exists that I could find; the author talks through the argument for about fifteen minutes on the LSE Centre for Economic Performance’s Innovation and Diffusion podcast (S3 E7, uploaded 25 June 2026), which is where the quotes below come from. No discussant. The interviewer says the paper is going to the Quarterly Journal of Economics and Juhász does not correct her.
Here is a fact about the Industrial Revolution that should bother you more than it does. Mechanized cotton spinning is invented in Britain in the late eighteenth century. It is enormously labor-intensive, which means it should want to run away to wherever labor is cheap, which by 1900 is everywhere that is not Britain. And yet in 1911 — a hundred and thirty years on — Britain and the United States between them still hold 61 percent of the world’s installed factory spindles, and the entire global periphery holds 22 percent. The technology is not secret. It is not patented anymore. The machines are for sale, from Platt Brothers of Oldham, to anyone with money. Wages in Osaka and Bombay and Shanghai are a fraction of wages in Lancashire. The technology stays in Lancashire.
The usual way to explain this is to say that adoption is hard: you need capital, institutions, complementary skills, a legal system, coal. All true. This paper says something narrower and stranger, which is that a large part of the problem was that the instruction manual was in English, and that in most of the world’s languages the words you would need to translate it into did not yet exist.
The friction is lexical, and the lexicon is a public good
Think about what it takes to build a spinning mill in 1880 if you are Japanese. Somewhere in Britain there is a book — “The American Cotton Spinner, and Managers’ and Carders’ Guide,” say, which the authors actually use — that tells you the dimensions of the building, how to set the gearing that distributes power through it, and how to run and maintain every machine on the floor. That book is the technology, in the sense that matters for you. You cannot read it.
So you translate it. Except you can’t, quite, because Japanese in 1860 has no word for telegraph, or steam engine, and you can only get so far writing English sounds in katakana, which produces transliteration rather than translation — a page of noises the reader has to memorize one by one. And even if you personally invent a word, the next translator invents a different one, and now there are two Japanese words for boiler and neither reader can read the other’s book. This is a plain coordination problem sitting underneath a plain public-good problem: someone has to fix the vocabulary for everyone, and then someone has to produce translations that anyone can copy the moment they exist. Markets are famously bad at both.
The Japanese state did both, deliberately, starting embarrassingly early. Within months of Perry’s ships arriving in Edo harbor in 1853, the shogunate stood up an institution whose name translates as the Institute of Barbarian Books, whose assignment was to build English–Japanese dictionaries so that technical translation would become possible. The solution its translators landed on was to build new Japanese jargon out of Chinese glyphs, which every literate Japanese person already knew, and which function roughly the way Greek and Latin roots function in English. Fukuzawa Yukichi coined denshin for “telegraph” in 1866 by gluing together the characters for electric and message. You do not know what a telegraph is the first time you read “electric message,” but you will remember it after someone tells you once, which is precisely the property that transliteration lacks.
The paper’s evidence that this worked is a graph of when Japanese words were first attested, taken from the Nihon Kokugo Daijiten. New-word creation in Japan runs at about a hundred a year for decades. It does not move when the Americans show up in 1853 with a working locomotive and a telegraph machine — which is worth sitting with, because it means physically showing a country your technology does nothing. It moves when the dictionaries come out.

Then the books follow the words. Between 1500 and 1860, Japanese translators managed eight Western technical books, total. By 1900 they had done 608. The stock of technical books in Japanese had been growing at 1.6 percent a year for two and a half centuries — a doubling time of forty-four years — and between 1870 and 1900 it grew at 8.8 percent, a doubling time of eight. Seventy-four percent of the identifiable translators were on the government payroll, which the authors treat, reasonably, as a lower bound.
The measurement problem, and a clever way around it
All of that is a country-level story, and country-level stories are where causal claims go to die. The move that makes this a paper rather than an essay is figuring out which industries should have benefited, without using any Japanese data to decide — because if you rank industries by what the Meiji government chose to translate, you have measured the government’s guesses about which sectors would succeed, and you have learned nothing.
So the authors build the ranking entirely out of British material. They digitize the synopses of every British patent from 1780 to 1852, out of Woodcroft’s Subject Matter Index. They hand-curate 460 English-language technical manuals covering the industries in their trade data. Then they ask, for each industry, how much the vocabulary of its manuals overlaps with the vocabulary of the patents — cosine similarity on TF-IDF-weighted unigram and bigram vectors, the workhorse measure:
Equation 1 in the paper. is the vectorized Woodcroft patent corpus, the vectorized manuals for industry , and the vocabulary size; the score is just the cosine of the angle between them.
They call it British Patent Relevance, and the nice thing about it is that it picks up input–output linkages for free without anyone having to code them: an industry that runs on steam engines will have manuals full of the bigram “steam engine,” and will score high, even though nobody patented that industry’s output. The ranking it produces is reassuringly sane. Textile yarn tops it at 0.146. Firewood and charcoal sits near the bottom at 0.007, along with lead, zinc, nickel, spices, and tea — things you dig up or pick, for which the First Industrial Revolution had nothing much to offer.

The outcome variable comes from what is, as far as the authors can tell, the first bilateral industry-level trade dataset for the nineteenth century — 37 regions, 93 industries, quinquennial from 1880 to 1910, stitched together from newly digitized Japanese and US records plus existing Belgian and Italian data, using the identity that j’s exports to i are i’s imports from j to fill in the non-reporting regions. The regression is a cross-section of industry growth rates with exporter fixed effects:
Equation 2 in the paper. is annualized export growth for region ’s industry from ~1880 to ~1910, an exporter fixed effect, and the two interactions let BPR have a different slope in Japan than in whatever comparison group is being tested.
What comes out
A one-standard-deviation higher BPR is worth 12 percentage points of extra annual export growth in Japan, and 1.2 percentage points of productivity growth in the Costinot-style version. Outside Japan the pooled coefficient is negative — minus 3 percentage points, significant at 1 percent. Which is the whole argument in one table: the same British knowledge that pulled Japanese industries up was pulling everyone else’s down, because everyone else was on the receiving end of it.

The specifications that matter are 5 through 7. If you pool the four languages that actually had technical literatures in 1870 — English, French, German, Italian — they get a positive coefficient too, 0.078, smaller than Japan’s, which is what you would expect from a country with less to catch up on. If instead you sort regions by income, or pull out Asia, nothing lines up: low-income regions have a large negative point estimate, Asia has a large negative one, and Japan looks nothing like either. So it is not distance-to-frontier and it is not geography. It is which languages had books.
The timing, which is the part that convinces
The cross-section still leaves open that some unobserved Japanese thing — Tokugawa literacy, culture, coal — explains everything. The answer is to run the same regression repeatedly, growing the window forward from 1875 five years at a time, and watch the coefficient.

Before Japan was technically literate, Japanese comparative advantage was moving away from the industries British technology had most transformed — exactly like the rest of the periphery, exactly like Asia. Then it flips. Not gradually: it flips right where the National Diet Library count goes from 706 technical books in Japanese in 1880 to 2,823 in 1890, which is the decade Japanese passes every European language except English and French.
That timing is very hard to get any other way. Japan opened to trade in 1858 and the coefficient does not turn positive for thirty-seven years. The Meiji Restoration is 1868 and it does not turn for twenty-seven. Tax reform, banking reform, the postal system, the telegraph, conscription, the constitution — nearly all of it was done by 1875, which is before the placebo window in which comparative advantage is still moving the wrong way. Whatever explains this pattern has to explain both a reversal and a date, and “institutions gradually improved” explains neither.
Here is what the counting looks like on both ends. In 1870 an Arabic reader had 71 technical books available in Arabic. Japanese was in the same neighborhood — indistinguishable from the rest of the periphery, and less codified than Spanish as late as 1880. By 1910 Japanese is third in the world.

The part I find hardest to argue with
The authors are careful to say codification was necessary and not sufficient, and they have an unusually good case for it, which is that they can show the same policy being deliberately copied and then failing. Park Chung Hee and Zhou Enlai both studied in Japan as young men, both came away Meiji admirers, and both, on taking power, launched state translation and technical-publishing programs. The book counts respond immediately: China goes from about a thousand technical volumes in 1949 to over thirty thousand by the early sixties, with the kink sitting exactly on Zhou’s ascension; Korea kinks in 1961 and again in 1982.

Korean incomes take off. Chinese incomes do not — Chinese GDP per capita in 1961 is barely above 1950, and growth does not arrive until after 1976, because in between there is the Great Leap Forward and the Cultural Revolution. The paper’s phrasing is that China is the exception proving the rule: totalitarianism can eliminate the benefits of codification. That is a slightly convenient reading of a case with a great many moving parts, and the authors do not pretend otherwise, but it does establish the thing a necessity claim needs, which is that you can have the books and still get nothing.
And the thing the whole paper is quietly about
The Meiji government did do the industrial policy you are picturing. It ran state model factories, subsidized favored firms, built railways, hired 2,400 foreign instructors for a total of 9,506 person-years, and sent the Iwakura Mission abroad. The literature on those policies is discouraging: Sussman and Yafeh conclude that “the great majority of the Meiji reforms” — the Bank of Japan, modern monetary policy, the constitution, parliamentary elections — “produced no quantitatively significant market response.” Twenty-five years after opening, eighty percent of Japanese exports were still primary products, and per capita growth was running at 0.6 percent.

What this paper nominates instead is the least glamorous item in the portfolio. As Juhász puts it, the policy “has a very different flavor… it’s much more horizontal in the sense that there’s no picking winners, there’s no picking sectors here in this specific policy. It’s really a public good provision” (22:16). No tariff, no champion, no credit allocation. A dictionary, a translation budget, and a school system built so that 40 percent of elementary class time went to scientific subjects — more than anywhere else at the time — such that by 1890, 90.6 percent of boys and 71.7 percent of girls were enrolled and could, in principle, read the manual.
It was not cheap, and the paper is good on where the money came from: the 1873 land tax reform, which Japanese economic historians call the single most important reform of the Restoration, and whose intellectual origin was a Tokugawa-era government translator who had rendered a book on economics into Japanese and worked out that a land tax raises more revenue with less distortion than an output tax. By 1884 it had given Japan an eight-to-one per-capita taxation advantage over China. Japan spent 11 percent of its 1880 budget on education alone; had China tried to copy only the education piece of the Meiji package, it would have had nothing left for anything else.
Which is a nicely closed loop, if you like that sort of thing. The translation program paid for itself with a tax that was itself imported by translation. Somebody read a book about how to collect money, in a language they could read, and used the money to buy everyone else books.
(Small honest note on the identification: the 1875–1880 placebo window has 71 industries in it and a coefficient of about −0.10 with a confidence interval running from roughly −0.20 to just below zero. It is significantly negative, but it is the noisiest point in Figure 10, and the reversal story leans on it. The tighter and more persuasive fact is the flip itself, which shows up in four independent patent corpora — later British patents, AI-summarized British patents, early US patents, late US patents — with coefficients between 0.111 and 0.121 every time.)