I Taught an AI to Read 26,000 MIT Theses. It Cost $4.
New technologies show up in a graduate thesis years before they show up in your life. I taught an AI to read 26,000 of them: $4, under an hour, ninety-two years of MIT’s archive.
The AI-reader that’s been going through Texas air permits started as a lab project at MIT six years ago, reading thesis archives. That project just got rebuilt. And you can talk to it.
The Texas work continues; this is the machine behind it.
New technologies show up in graduate theses years before they show up in your life. Long before the product, the pilot program, the utility interconnection queue, or the public fight at a zoning hearing, somebody sat in a library and wrote two hundred careful pages about an idea most people hadn’t heard of yet.
In 2010, before data centers became the public face of the AI buildout, MIT theses were already treating them as physical infrastructure. Not as an abstract computing topic, but as power loads, cooling systems, airflow models, operating cost, and energy efficiency. One thesis piloted low-cost efficiency improvements at a Raytheon data center in Garland, Texas, reporting a 23% reduction in electricity use. Another modeled above-plenum airflow with CFD to understand cooling losses above a raised-floor data center.
The market would later call this a real estate, energy, or permitting problem. The Archive shows it arriving first as a building systems problem.
That gap, between when an idea enters the archive and when it enters the world, turns out to be measurable.
Back in 2020, we built a tool at MIT’s Real Estate Innovation Lab to trace it. It took a lab, a team, and an industry partnership with JLL. This month I rebuilt it: $4 in compute, under an hour.
You can explore Attention Archive here, or just talk to the LLM Version.
This essay is how it works and what it found.
The $4 Rebuild
The $4 isn’t a gimmick. It’s a finding before the findings.
The hard part of the 2020 project wasn’t computers. It was reading time. What changed is that reading became a thing you can buy by the token. The whole archive, every thesis individually, for less than a typical Starbucks order. Any question that was ever too expensive to ask of an archive is now worth asking.
So here’s what asking looks like.
Attention Archive reads 26,562 MIT theses across Civil and Environmental Engineering, Architecture, EECS, Urban Studies and Planning, and Management. Not just keyword search. An AI reader goes through each thesis and records what it is actually about: the technologies discussed, the methods used, the geography, the themes, and the one thing that turned out to matter most: how the thesis talks about each technology.
That “how” gets scored on a five-stage scale, from the language of wondering to the language of taking for granted:
speculative → pilot → evaluative → critical → infrastructural
A speculative thesis asks whether something will matter. A pilot thesis builds with it and reports back. An evaluative thesis measures whether it worked. A critical thesis asks what it breaks. An infrastructural technology barely gets discussed at all. No thesis argues for concrete. Concrete is just there.
Score decades of writing this way and every technology traces a path. You can watch attention move: sustainability overtaking everything after 2000, computing going from footnote to method, AI going vertical, and data centers moving from equipment rooms into the built-environment conversation.
Keyword vs. Intention
Here is the uncomfortable part.
Most of what anyone believes about research trends, including what I believed in 2020, comes from dashboards that count keywords. Trend reports, horizon scans, “state of the field” charts: pattern matching, mostly, all the way down.
Pattern matching can be wrong in a very specific way, and now we can measure it.
Keyword mode works the way most dashboards work. Here is the actual rule for the “Data, GIS & Computing” theme, straight from the config, because methods should be checkable:
/\bgis\b|geographic information|\bdata\b|comput|simulation|\bmodel|algorithm|machine learning|sensor|digital|software/i
AI-read mode asks a different question: not does this thesis mention data, but is this thesis actually about data, GIS, or computing?
Same archive, same year, very different answers. In 2024, keyword matching tags 64% of built-environment theses as computing research. The AI reader says 15%. Skim the disagreements and the pattern is obvious: a pharma logistics thesis “uses data,” a housing policy thesis “builds a model,” a planning thesis cites simulation. Keyword counts all of it as computing research. On actual reading, only about one in five of those keyword hits turn out to be about computing at all.
That’s the difference between keyword and intention. A word appearing in a document tells you the author touched a subject. Only reading tells you whether they meant it.
And the gap itself is the story. After 2000, the keyword line climbs toward 70%, not because everyone switched to computing research, but because data, models, and simulation became how all research gets done. Computing is both a subject and a tool, and those are not the same thing. The distinction matters because it is exactly how technologies become infrastructure: first they are strange enough to be studied directly, then they become useful, then they become assumed.
The old instrument could not tell the difference between a field’s topic and a field’s toolkit. Reading can.
If your dashboards run on keyword matching (and most of the world’s do), that 64-vs-15 gap is worth sitting with.
The Data Center Question
This is why the data-center example matters.
If you ask when data centers first appear, the archive does not just return a trend line. It returns receipts. The early built-environment examples are about energy efficiency, facilities management, airflow, cooling, and operations. By 2025, data-center attention climbs to its highest level yet in the tracked series. The archive makes visible a shift that now defines the market: computation becomes real estate when it starts needing land, power, cooling, capital, and permission.
That crossing, from one discipline’s equipment room to another discipline’s land fight, took about fifteen years. And the archive caught a second crossing that is happening right now.
Data centers made that crossing between 2010 and 2025. Generative AI is making it now, and this time the instrument is watching in real time.
That is also why this connects back to Texas air permits. The permit archive is not a separate project. It is downstream of the same question: when does a technology stop being software and start becoming infrastructure?
Ahead of the Patents
One more early result, flagged as early.
I ran the same AI reader on a pilot set of built-environment patents, using the same schema, and lined up the attention curves. For the first overlap set of technologies, thesis attention peaks roughly sixteen years before patent mention volume. Median, four technologies. A pilot, not a law.
But the direction is interesting. Economists Comin and Hobijn estimated that technologies can take decades to fully diffuse into use. The MIT corpus sits at the front of that lag. What this means is that the academic archive is not just an old record of what already happened. It is a leading indicator of what is coming, sitting in public, machine-readable, and mostly unread.
Kick the Tires
Everything above is checkable, and that is deliberate. The keyword baseline is published. Technologies enter the lifecycle chart by rule rather than taste. Small samples are marked. Answers include thesis links.
And as of this week, you don’t have to take my word for it. Ask the archive directly.
Ask when data centers became an energy problem. Ask when GIS first appeared, and in whose thesis. Ask where keyword search and AI-read disagree. The system answers with numbers, sample sizes, charts, and links to the theses themselves. It also tells you when the archive can’t answer.
If you want the extraction schema and the keyword-vs-AI-read comparison notebook, reply to this email or comment “method” and I’ll send it.
What should this Reader read next? Patents at scale are the obvious candidate. Funding rounds and trade press would test whether the lead-lag holds downstream. A city’s planning archive is already in the works.
If you have access to a corpus and want to see it read, my DMs and inbox are open.
Attention Graph: archive.infrastructure-research.com
LLM version: ai.archive.infrastructure-research.com






