Image for The Researcher's Path Part 11: Formalization
Technology Jun 24, 2026 • 15 min read

The Researcher's Path Part 11: Formalization (Turning the Narrative into Notation)

Part 11 of The Researcher's Path. How to translate a working framework into language precise enough for strangers to verify. The hidden sephirah, Da'at, is where story becomes structure.

Share:
Lee Foropoulos

Lee Foropoulos

15 min read

Continue where you left off?
Text size:

Contents

The Researcher's Path: A 13-Part Series

Part 1: Environment SetupPart 2: AI ConditioningPart 3: Literature SurveyPart 4: Root QuestionPart 5: ClassificationPart 6: StructurePart 7: ExpansionPart 8: Critical AnalysisPart 9: IntegrationPart 10: Force MappingPart 11: FormalizationPart 12: Pattern RecognitionPart 13: Publication

There's a moment every researcher hits where the work is true and still can't leave. You can explain it at a whiteboard, you can hold the whole thing in your head, you can watch your framework produce correct predictions on new data. And if someone asks "where can I read this," you realize you don't have an answer that would let them reconstruct what you know. The work is real. The work is not transmissible.

A handwritten notebook beside neatly typed manuscript pages. The visible moment a private framework is being turned into something a stranger can read
Formalization is the work of dragging what you know in your head onto a page where someone you've never met can pick it up and use it.

That gap is what Part 11 is about.

A framework that only works in its inventor's head isn't a framework yet. It's a talent.

"The work is real. The work is not transmissible."

The sephirah at this stage is Da'at. Knowledge. It doesn't appear on all versions of the Tree of Life; it's the hidden sephirah, the one some traditions draw and others omit. Da'at sits at the gap between the three supernal spheres above and the seven below. Its job is bridging. In classical Kabbalah it's where the abstract flashes of Kether, Chokmah, and Binah get articulated in a way the lower sephiroth can use. In research, formalization is the same motion: you take what you know pre-verbally and drag it into a form that doesn't lose anything when it leaves you.

Why formalization is harder than it looks

The common mistake is to treat formalization as note-taking. As if the framework already exists and all you have to do is write it down. It isn't note-taking. It's a translation. And every translation costs you something.

What you know at the end of Part 10 is a working intuition. You've watched it make predictions. You've seen it survive foils. Parts of it are crisp and parts of it are waving hands. The waving-hands parts are the ones that matter now, because a reader who can't see your hands has only the paper. If the paper hand-waves, they'll either reject the framework or, worse, accept it for the wrong reasons and build on a foundation you can't vouch for.

≥1
load-bearing sentence per median paper that the author would admit, in private, they couldn't defend under questioning. Formalization is the work of finding those sentences before the reviewers do

The three layers

A close-up of research notation. Equations, terminology definitions, and derivations laid out in three visible blocks on a page
Vocabulary on top, structural claim in the middle, predictions at the bottom. Each layer constrains the next.

A formalized framework has three layers, and they have to be written in that order because each one constrains the next.

Layer one is the vocabulary. Every term you use has to mean one thing. If "force" appears in your paper, it has exactly one technical meaning there, and that meaning is defined on first use. This sounds trivial until you try it. You'll find words in your own draft that you've been using to mean two related-but-different things. One meaning when you're describing the mechanism, another when you're describing the prediction. Collapse them, or rename one.

Layer two is the structural claim. In prose, the framework reads as a story. In the formal layer, it reads as a set of axioms and derivations. You don't need the derivations to be in the Greek-letter, quantified style of physics. For most fields that's overkill and for some it's wrong. But you do need them to be in the form "this assumption; which implies this; which implies this." Anyone reading the formal layer should be able to point at the exact step where they disagree, rather than gesturing at the whole thing.

Layer three is the predictions. This is the force map from Part 10, rewritten in the vocabulary of layer one and citing the structural claims of layer two. The predictions should be mechanical consequences of the structure. Not additional assumptions, but things that have to be true if the structure is true. If a prediction doesn't follow from the structure, it means you have a hidden assumption you haven't made explicit; back up and fix layer two.

A reader who can see where the rigor lives and where the narrative lives is a reader who can trust the paper.

The test for layer two

After you've written the structure, try to derive a prediction you haven't used yet. If you can, you're doing formalization right. The framework is productive. If you can only recover the predictions you already knew about, the framework is a summary, not a theory. Keep working.

Language choice

Pages comparing a differential equation, a probability graph, and a category-theoretic diagram side by side. Three formal languages competing for the same idea
Three formal languages, one underlying claim. The right choice is the one that makes the structural claims easiest to check and the derivations hardest to fudge.

The most common error in formalization is reaching for the most prestigious language rather than the most faithful one. Not every framework wants to be expressed in differential equations. Not every framework wants to be expressed as a probabilistic graphical model. The right formal language is the one that makes the structural claims easiest to check and the derivations hardest to fudge.

For a framework about causation in a small, well-controlled system, algebra on finite structures may be the right language. For a framework about emergent behavior in a large system, information-theoretic measures may be the right language. For a framework about relationships across domains, category theory, oddly, is sometimes the cleanest option. The wrong answer is to pick the language your field defaults to just because your field defaults to it.

1
translation layer the reader is always owed: the bridge between the formal object and the everyday object it represents. Skip this and the formalism floats free and impresses no one

Whatever you pick, you owe the reader one piece of plumbing: a translation layer between the formal object and the everyday object it's meant to represent. If your formal object is a graph and your everyday object is a pattern of citations, tell the reader which nodes are papers, which edges are citations, and which graph-theoretic property corresponds to which real-world concern. Otherwise the formal work floats free and impresses no one.

The temptation to over-formalize

Once you discover how much the formalism helps. How it tightens the thinking, how it catches contradictions you didn't see, how it makes the reader unable to evade your claims. You'll be tempted to push the formalism into every corner of the framework. Don't. There are parts of most good research that are genuinely narrative: historical context, motivation, interpretation, implications. Formalizing those strips out the human reasoning and replaces it with a false precision.

Formalize the structural claims. Formalize the predictions. Leave the motivation and the interpretation in prose. A reader who can see where the rigor lives and where the narrative lives is a reader who can trust the paper. A reader who can't distinguish the two starts to suspect the rigor is decorative. They're often right to.

Keep a narrative sibling

Every formal section should have a prose sibling that says the same thing in plain language, just before or after. The sibling isn't for the reader who can't read the formalism. It's for you. If you can't write the sibling, you don't actually understand your own formalism yet, and the reader won't either.

The Da'at problem

A bare workbench with a pencil, a printed paper, and a single equation chalked on a slate. The unromantic crossing where private knowledge is forced into public form
Da'at is the bridge. The framework feels real before it crosses, and only real research after.

Da'at is controversial in Kabbalah for a reason. Some traditions refuse to count it as a sephirah at all, because knowledge that can't be transmitted is just a private experience, and knowledge that can be transmitted has already been translated into one of the other sephiroth's languages. Logic, compassion, rigor. The bridge doesn't deserve its own sphere.

That controversy is exactly the tension of formalization. The framework, in its pre-formal state, feels like it deserves a sphere. It feels like it's real knowledge. And it won't be real knowledge until it's been pushed into one of the existing formal languages and stripped of everything the languages can't carry. Some of what's lost in that stripping is genuinely important. The intuition, the taste, the judgment that guided the research. Some of what's lost was never there to begin with, and its loss was what formalization was for.

Afterward
is when you find out which parts of the intuition were real knowledge and which were comforting noise. There's no way to know before you formalize

You don't know which is which until afterward.

Formalization doesn't preserve everything you know. It preserves only what you can show.

Coming up for air

After Part 11 the framework is transmissible. Someone else could, in principle, take your formalism and run the same tests you ran. That's a massive step. It's also not enough. A framework that's transmissible to one stranger isn't necessarily transmissible to the field, and the predictions that worked once aren't necessarily repeatable across time and samples. Part 12, Pattern Recognition, is where you stop trusting your own data.

Your formalization checklist 0/5
How was this article?

Share

Link copied to clipboard!

You Might Also Like

Lee Foropoulos

Lee Foropoulos

Business Development Lead at Lookatmedia, fractional executive, and founder of gotHABITS.

🔔

Never Miss a Post

Get notified when new articles are published. No email required.

You will see a banner on the site when a new post is published, plus a browser notification if you allow it.

Browser notifications only. No spam, no email.

0 / 0