Thursday, September 17, 2026

Natural Language Is Not Going Away

Quite a few weeks ago, my friend Simone Severini sent me a book he wrote with Andrea Borghini, Science, Ltd.: Notes on the Research Enterprise in the Age of Machines. He told me, rather modestly, that they had quoted me in it.

It turns out that they quoted me quite a lot.

The conversations behind the book go back to May 2023, although the message Simone sent me about the finished book arrived in August 2025. Reading the book now was an interesting exercise, because I discovered that I still believe everything that I told him then—and that much of what I have been working on in the intervening years has gone in precisely that direction.

The book is about a large question: how should the scientific enterprise change when machines become capable not simply of calculating things for scientists, but of reading, organizing and connecting scientific knowledge at scales that humans cannot?

One of the problems Simone and Andrea discuss is particularly acute in mathematics. We have millions of papers. We have more mathematics than any individual could possibly read. But the relationships between mathematical ideas are still mostly buried inside those papers, expressed in the language mathematicians use to talk to one another.

This was the context in which they asked me about formalization.

I apparently said:

“My goal is to try to overcome this lack not of data, but of annotations on data when working in mathematics, so that we can explore the scientific literature and understand texts in depth. And we should try to extract as much information as possible with natural language before proceeding with formalization.”

I still think this is right. (annotations here is an euphemism for semantics)

There is a very understandable temptation, when thinking about mathematics and machines, to conclude that the solution is to formalize everything. Formal mathematics is great, and formal proof assistants are becoming extraordinarily powerful.  But formalization requires choices: a formal system, a representation, a particular piece of software, conventions about what is made explicit and what isn't.

Mathematicians themselves don't normally communicate that way. We talk, write papers, draw diagrams, invent terminology, abuse notation, leave things implicit and rely on enormous amounts of shared mathematical culture.

So I also told them:

“When we collaborate, we speak in our language. Moreover, when you decide to formalize something, you have to make a technical choice and adopt a tool that, all of a sudden, might become obsolete as soon as a better formalization is discovered. See programming languages. However, it is unlikely that the use of natural language will become obsolete: it is an excellent tool that has evolved over thousands of years.”

Here, for once in my life, I underestimated a number. (Brazilians are known for exaggerating them!)

Language has been with us for considerably longer than “thousands of years.” Exactly how much longer is a fascinating and disputed question, because spoken language leaves no fossils. Estimates for fully developed human language often reach well beyond 100,000 years, and some researchers argue for a much older origin still.

So perhaps I should have said: natural language is an extraordinarily successful technology that humans have been developing for at least tens of thousands, and very possibly hundreds of thousands, of years. It would be surprising if the arrival of proof assistants and large language models suddenly made it obsolete.

What I wanted instead—and still want—is a good interface between these two worlds.

In the years since those conversations, this has increasingly become, for me, the idea of Network Mathematics: representing mathematical knowledge not as a pile of documents and not as a single enormous formal library, but as an extended knowledge graph connecting the different ways in which mathematics exists.

Some of our work since then has been about trying to build pieces of precisely this bridge. MathGloss tries to identify and align mathematical concepts across resources written for humans and resources intended for machines; more recently, our work on extracting mathematical relations asks whether language models can recover some of the relationships that mathematicians leave implicit in mathematical prose. Concepts are hard enough to identify; relations between them are harder still. But extracting both may give us enough overlapping evidence to begin reconstructing the network that is already there in the literature. 

Definitions, concepts, theorems, examples, papers, informal explanations, formal statements, proofs, databases and libraries should be connected. A mathematical concept appearing in ordinary mathematical English should be able to point toward its occurrence in Wikipedia or Wikidata, toward related concepts, toward its use in papers, and eventually toward one or several formalizations of it. 

The point is not to replace informal mathematics by formal mathematics, nor formal mathematics by informal mathematics.

The point is to build the bridges.

 This is why I was particularly pleased to discover that Simone and Andrea return to our conversation later in the book, when they start imagining “science in a graph.” They ask what the vertices of such a graph should be, what language should describe them, and how the connections between them should be represented. Existing scientific databases give us the documents, but most of the interesting conceptual relationships are still hidden inside those documents and recoverable only by readers who already understand the field.

Yes. Exactly.

This is also why the recent explosion of language models seems important to me. If machines are becoming much better at dealing with the language scientists actually use, perhaps we don't have to choose between the enormous human inheritance encoded in natural language and the precision and computational possibilities of formal systems.

We can try to connect them.

Simone and Andrea's book goes much further than mathematics. It asks how publishing, peer review, funding, scientific databases and ultimately the institutions of science might have to change if scientific knowledge becomes genuinely machine-readable and machine-navigable. It is also appropriately suspicious of the idea that simply making science faster automatically makes science better.

But I was delighted to find my 2023 self appearing in their argument.

Especially because, three years later, I would give them essentially the same answer. Despite all the proposals and papers rejected on that. I am stubborn.

Except for the number.

No comments:

Post a Comment