Axis · entirely unproven
Worth reading not because it's correct — because it might be.
Three questions, and none of them are about computers.
What is the sky saying? What are the animals saying? And what are we saying to each other, underneath the words we actually use?
Each of those is the same problem wearing different clothes. Something out there is producing structure. We can measure it, in extraordinary detail. And we still cannot say what it means.
We are not proposing to solve any of them. Every one of those questions already has a field of its own, with its own instruments and its own vocabulary, and most of them are good at it. What none of them have is a way to tell each other what they found. Someone listening to whales and someone watching a fault line are doing the same work, and will never read each other.
This page argues for one thin thing in common — so that a pattern found in one place can be recognised in another, and checked. That is a smaller claim than it first appears, and a much harder one to dismiss.
signal ──▶ CORE IR ──▶ meaning
Once you see that shape, the question stops being what is this language for and becomes what counts as a signal.
whale codas seismic precursors stellar light curves
bee waggle dances volcanic tremor gravitational waves
neuron spike trains ocean current shifts solar magnetic flux
birdsong dialects plate strain radio bursts
protein folding grid load oscillations market microstructure
gene expression traffic flow climate teleconnections
immune response crystal growth cell differentiation
Every one of those is the same problem wearing different clothes: something is producing structure, we can measure it, and we cannot say what it means.
We currently build a separate field, a separate toolchain and a separate vocabulary for each. Nothing a seismologist learns about finding structure in noise is expressible to someone studying birdsong — not because the problems differ, but because there is nowhere shared to put an answer.
The work is already being done. It's being done in parallel, by people who will never read each other.
bioacoustics ── HMMs, unsupervised clustering over calls
seismology ── precursor detection, template matching
bioinformatics ── Gene Ontology, motif discovery, alignment
NLP ── embeddings, topic models
ecology ── early-warning indicators, regime-shift detection
finance ── factor models, change-point detection
neuroscience ── state-space models, energy landscapes
epidemiology ── phase detection, R-t estimation
HCI ── dwell time, trajectory curvature, paradata
Every one of those is semantic analysis: signal in, structure out, meaning attempted. Every one has its own toolchain, its own vocabulary, its own conferences and its own journals.
A regime-shift method validated in ecology cannot be expressed to a seismologist. Not because the problems are unrelated — because there is no carrier. And the cost is invisible, because nobody ever experiences it. You simply never find out about the method in the field next door.
The architectural pattern for fixing this is old and well understood.
many domains, each with its own methods
┌──────┬──────┬──────┬──────┬──────┬──────┐
│ bio │ eq │neuro │ eco │ econ │ ... │
└───┬──┴───┬──┴───┬──┴───┬──┴───┬──┴───┬──┘
└──────┴──────┴──┬───┴──────┴──────┘
▼
╔══════════════╗
║ CORE IR ║ one thin invariant.
╚══════════════╝ everything passes through it.
▲
┌──────┬──────┬──┴───┬──────┬──────┬──────┐
│ ML │stats │ dyn │ info │ topo │ ... │
└──────┴──────┴──────┴──────┴──────┴──────┘
many methods, each domain-agnostic
TCP/IP (Cerf & Kahn, 1974) solved nothing about applications and nothing about link layers. It gave incompatible networks one thin thing in common, and everything above and below it became interchangeable. That is the entire proposal here, applied to meaning rather than packets.
Unification of this kind has a history, and it is not uniformly encouraging — which makes the pattern in the failures the most useful thing on this page.
| Attempt | Outcome | Why |
|---|---|---|
| SI units and metrology | worked | a physical artifact, later a defined constant |
| Positional numerals | worked | a notation you compute in, not one you talk about |
| TCP/IP (Cerf & Kahn, 1974) | worked | a wire format, machine-checkable |
| Unicode | worked | a code point is an identity |
| Gene Ontology (Gene Ontology Consortium, 2000) | worked, narrowly | annotations became comparable across labs and species |
| Dimensionless numbers (Buckingham, 1914) | worked | Reynolds number transfers between systems sharing no substance |
| General Systems Theory (von Bertalanffy, 1968) | failed | vocabulary only |
| Cybernetics (Ashby, 1956) | faded | vocabulary only |
| Semantic Web / RDF | stalled | an artifact, but authored alongside the real thing — so it drifted |
The variable is whether there was a mechanical carrier, or only a shared vocabulary.
Von Bertalanffy (1968) was right about isomorphisms across systems. He left behind a vocabulary and nothing anyone could run, and the whole programme dissolved into metaphor within a generation. That is the failure mode to name out loud, because it is the one this proposal is closest to.
The difference claimed here is narrow and specific: a hash can be checked and a word cannot. Two researchers using the same word may mean different things and have no way to find out. Two structures with the same identity are the same structure, and that is not a matter of opinion.
The one to put in front of a sceptic is dimensional analysis (Buckingham, 1914). A model ship in a test tank predicts the behaviour of a real ship because the two share a dimensionless structure — not a substance, not a scale, not a material. Cross-domain transport by structural identity already works, and has worked for a century, in the one place someone bothered to formalise it.
This gap is not newly noticed. It was formally identified, named, and then left alone for seventy-seven years.
Shannon's 1948 paper (Shannon, 1948) states plainly that the semantic aspects of communication are irrelevant to the engineering problem. That wasn't an oversight — it was the scoping decision that made the theory work. Weaver's introduction the following year (Weaver, 1949) then set out the three levels the field would face:
LEVEL A technical how accurately can symbols be transmitted?
── solved. 1948. it built everything.
LEVEL B semantic how precisely do the transmitted symbols
convey the desired meaning?
── named. still open.
LEVEL C effectiveness how effectively does the received meaning
affect conduct in the desired way?
── named. still open.
Level A got a complete mathematical theory and the entire communications industry. Level B got a name.
This page is a proposal for Level B, and that framing is deliberate: it isn't a new field, it's an old and explicitly identified one (Shannon, 1948; Weaver, 1949) that has been waiting for tooling that didn't exist in 1949.
The claim, once, plainly: every field has already built its own semantic analysis, and none of them can talk to each other. This is not a proposal for better analysis. It is a proposal for the thing they would need in common for their existing results to become each other's evidence.
Machine learning is now extraordinary at signal → something. AlphaFold. Earthquake nowcasting. Exoplanet detection. Protein language models. Sperm whale codas (Sharma et al., 2024).
It finds the structure. Then look at where the finding lands:
a pattern is discovered ──▶ a 4,096-dimensional embedding
├── opaque nobody can read it
├── model-specific another model's is incomparable
├── mortal retrain, and it's gone
├── unfalsifiable there's no claim to argue with
└── trapped it can't leave its domain
The pattern is real and the answer has nowhere to live. We have built the most powerful structure-detectors in history and we write their findings into a format that cannot be compared, cannot be checked, and dies with the checkpoint.
Core IR is a place to put it: explicit, identified, falsifiable — and surface-independent, so the answer isn't in seismology's language or biology's language or English.
If a discovered structure has an identity, structures from unrelated domains become mechanically comparable.
whale coda structure ──▶ CORE IR ──▶ hash 7f3a…
seismic precursor pattern ──▶ CORE IR ──▶ hash 7f3a…
▲
not a metaphor.
the same structure.
Right now, noticing that two fields share a deep structure is a rare act of individual genius that cannot be verified. Someone sees that predator-prey dynamics and combustion share a shape, writes it up, and the field argues about whether the analogy is real for twenty years.
Make it a query and analogy becomes computable. Cross-domain structural identity stops being insight and becomes a lookup against every pattern anyone has ever recorded.
A search engine for isomorphisms. The thing a synthesist does by hand, turned into infrastructure anyone can run.
AI finds structure in domain A
│
▼
recorded as Core IR — identified, comparable, permanent
│
├──▶ does anything, in any field, share this identity?
│ │
│ ▼
│ domain B has it — and in domain B, we already know what it MEANS
│ │
│ ▼
└──▶ that is a hypothesis about domain A. testable. this week.
Meaning transfers across the gap. Not because the domains resemble each other — because the structure is the same object, and one side of it was already grounded.
A machine for generating real hypotheses, where the expensive half — finding the pattern — is the half AI already does well.
The one that makes people laugh, and then stop laughing.
Project CETI have the structure and explicitly not the meaning: 9,000 sperm whale codas, context-sensitive combinatorial structure, published in Nature Communications in 2024 (Sharma et al., 2024) — with the lead author saying plainly that they don't yet know what the whales are saying, and that behavioural context is the next step.
Structure solved. Grounding open. That is exactly the gap this shape addresses.
And the field is trying to decode animal signal into English, which is the wrong direction — English is a surface too, and every human concept and category boundary rides along with it for free.
THE FIELD TODAY
whale codas ────────────────────▶ English
THE OTHER DIRECTION
whale codas ──▶ ┐
├──▶ CORE IR ◀── a third thing. neither species' surface.
human speech ──▶ ┘
then ask: do these ever land on the same node?
A negative result is still a finding — it tells you how alien the meaning space is.
And two-way communication, which is what actually matters, needs generation. One meaning, one rendering bridge per species, is the same object as a Rust bridge and a JavaScript bridge. Built for something else entirely, and already the right shape.
The first thing you should be able to say to a whale is "I don't understand." Every interspecies protocol starts with nouns — fish, danger, hello. But the message that makes a channel robust is the one that handles failure. UNKNOWN as a real state (Kleene, 1952) means that didn't parse, say it differently is expressible on both sides. That's how you bootstrap a shared vocabulary with no Rosetta Stone: not by guessing right, but by having a defined way to be wrong.
Plate strain. Volcanic tremor (cf. Haken, 1977, on the slaving principle near a transition). Ocean circulation. Stellar light curves. Solar magnetic flux. Radio bursts.
All signal. All structure. All currently interpreted by a separate community with its own vocabulary, and none of it comparable to anything outside itself.
The specific question worth asking: does the structure preceding a rupture look like the structure preceding any other commitment? Not poetically — by identity. Attractor-reconstruction methods (Takens, 1981) and convergent-cross-mapping causal detection (Sugihara et al., 2012) already do exactly this kind of state-space comparison within single domains — extending it across domains is the whole ask. If it does, seismology inherits everything any other field knows about that shape.
The only meaning substrate we know works, and we study it by measuring signals and guessing.
Spike trains, oscillations (Beggs & Plenz, 2003, on neuronal avalanches as scale-free cascades), connectivity, blood flow. We can measure all of it and read none of it.
The inversion worth noticing: this isn't an attempt to build something brain-like. It's an attempt to do explicitly what a brain does implicitly — hold meaning, identify it, relate it. So a brain recording ought to lower into it. If it doesn't, that's informative about the substrate — which would be the first falsifiable thing anyone has said about a meaning representation in a long time.
We diagnose states. Healthy, diseased, stage II. But disease is a transition — a cell holding one context and then committing to another, in the sense of Waddington's (1957) epigenetic landscape. The commitment is the event. The state is the aftermath, and by the time you can classify it you are already late.
TODAY measure the state ──▶ classify ──▶ "you have X"
HERE measure the transition itself
──▶ the structure of a system approaching commitment
──▶ and that structure may not be tissue-specific,
species-specific, or even biological
And the genome underneath it is a program nobody can read that we've been decoding by aligning strings. Expression is context-dependent interpretation, not lookup — same sequence, different context, different meaning, which is the definition of a surface. The sequence is the surface. The meaning isn't in it.
AlphaFold cracked structure from sequence and it is magnificent, and it still doesn't tell you what anything means. Same gap as the whale gap.
The largest signal source on the planet is people talking, and we analyse it by counting keywords and collapsing meaning into a single number between −1 and +1.
Sentiment analysis is the recovery industry in its purest form. An utterance carries structure, intent, hedging, commitment, contradiction and reference — and we throw all of it away, keep a scalar, and then build dashboards on the scalar.
Lower an utterance into Core IR and point as many analyser bridges at it as you like. One lowering, many readings, none of them re-parsing anything:
a conversation
│
▼
CORE IR ── what was actually meant: claims, hedges,
│ commitments, conditions, contradictions
│
┌────┼──────┬──────────┬───────────┬────────────┐
▼ ▼ ▼ ▼ ▼ ▼
sales negotiation clinical team health teaching support
│ │ │ │ │ │
└────────┴──────────┴─────┬──────┴────────────┴─────────┘
▼
five readings of ONE meaning.
where they disagree is itself information.
Today each of those is a separate product with its own NLP pipeline, its own model, and its own way of being subtly wrong. They are all re-deriving meaning from text that had it.
The replication crisis is usually framed as p-hacking and small samples. There's a deeper layer: a construct is often defined by the instrument that measures it — the jingle and jangle fallacies named a century ago (Thorndike, 1904; Kelley, 1927). Anxiety is what the questionnaire measures. Two labs use the same word for different things, or different words for the same thing, and there is no mechanism to tell which.
A construct with an identity separate from its instrument would make "did we study the same thing" a lookup instead of a twenty-year argument.
We assess collective state with GDP, unemployment, polls and surveys. Every one is a lagging proxy, sampled rarely, at enormous cost, and usually telling you about a transition that already happened.
And it's the same recurring shape as everything else on this page:
measuring the STATE ──▶ always late, always a proxy
measuring the TRANSITION ──▶ the structure of a system
approaching a commitment
Which is the disease argument, and the ecosystem argument, and the seizure argument. Critical slowing down has already been reported in financial systems and social ones (Scheffer et al., 2009) — so a community approaching a transition may carry the same second-degree signature as a cell approaching commitment.
That's either a profound claim or a meaningless one, and the point of having a substrate is that it becomes possible to find out which.
Some documents are written to be read. Others are written to be survived — long, flat, deliberately unmemorable, with the thing that matters buried on page forty in a subordinate clause.
Contracts. Terms of service. Loan agreements. Insurance policies. Legislation. Tender responses. Length and tedium are the concealment mechanism, and it works because a human reader's attention is finite and the drafter's isn't.
Reduce a document to meaning and tedium stops protecting anything:
what a CLAUSE means, not what it says
│
├── what obligations does this create, on whom, when?
├── what conditions are attached, and to what?
├── which clauses interact? (the sting is almost always
│ an interaction between two distant clauses)
├── what's asymmetric between the parties?
└── what changed from the last version — in MEANING,
not in wording?
That last one is the practical killer. A redline shows you every word that moved. It cannot tell you which of those changes altered an obligation and which were cosmetic. Abstract Meaning Representation (Banarescu et al., 2013) already demonstrates this at sentence scale: sentences that mean the same thing produce the same graph. Same meaning, same hash — a reworded clause that means the same thing is provably harmless, and everything that isn't harmless is what's left.
Across languages, the same mechanism answers a question nobody can currently answer. EU legislation exists in 24 languages, every one of them equally authoritative. Multilateral treaties are signed in several authentic texts. International contracts are executed in two. Where those texts diverge in meaning — and they do — it is discovered by litigation, years later.
Lower each authentic text independently and compare identities. Convergence is a check. Divergence is a finding, located at the clause.
Here's where it stops being document analysis and becomes a live instrument.
This is not speculative. Nine real hostage negotiations were divided into time stages and the dialogue of police negotiators and hostage takers analysed across eighteen linguistic categories (Taylor & Thomas, 2008). Successful negotiations showed higher linguistic style matching — the degree to which two parties coordinate their word use — than unsuccessful ones.
And the shape of the failure is the interesting part. Unsuccessful negotiations weren't characterised by a low level of coordination. They were characterised by dramatic fluctuation — negotiators unable to hold the steady rapport that the successful ones maintained. Successful dialogue showed more coordinated turn-taking, reciprocated positive affect, and a focus on the present rather than the past.
Read that again in the language of this page:
the level of coupling ──▶ first degree
the STABILITY of coupling ──▶ second degree ← this is what predicted outcome
variance rising before it
comes apart ──▶ the same signature as every
other approaching transition
Instability in the coupling predicted failure. That is critical slowing down's cousin, in a conversation, measured in a police negotiation, and published in 2008 (Taylor & Thomas, 2008) — by people with no interest whatsoever in proving anything on this page.
There's structure underneath it too: a cylindrical model of crisis negotiation (Taylor, 2002) describes three levels of interaction — avoidance, distribution, integrative — across three thematic styles: identity, instrumental, relational. That's a structured meaning space for a conversation, already mapped by the field. It just has nowhere to live where anyone outside crisis negotiation can use it.
The application that matters most is the one where getting it wrong is worst.
People in mental health crisis are met by responders who have seconds, no history with them, and no instrument beyond their own judgement. The outcomes are frequently terrible and very well documented.
What the above suggests is modest and useful: an indicator of whether the current approach is working, in near real time, against this person's own baseline. Not a diagnosis. Not a risk score. Not a recommendation about a human being. A read on coupling — is this conversation converging or coming apart — which is exactly what the negotiation literature (Taylor & Thomas, 2008) says predicts the outcome, and exactly what a stressed responder loses track of first.
And the same principle would let the field learn. Every de-escalation, encoded as structure rather than as a written report, becomes comparable to every other one. What actually works becomes a question with an answer, instead of accumulated craft that dies when an officer retires.
Four things must be said plainly, because this is the highest-stakes idea on the page:
Here's the one that should be obvious and isn't: the process of arriving at an answer carries meaning the answer doesn't.
A survey records that someone ticked agree. It discards that they took eleven seconds, changed the selection three times, scrolled back to re-read the previous question, and answered the next four in under a second each. The tick is kept. Everything that tells you what the tick was worth is thrown away.
WHAT WE KEEP WHAT WE DISCARD
the answer ◀──── time to first keystroke
pauses mid-sentence
deletions and rewrites
revisions of a selection
dwell on one item
scrolling back to re-read
order of completion
the point at which they gave up
Which is the same move as everywhere else on this page: we keep the state and discard the transition. The answer is the state. The hesitation is the transition, and it's where the information about confidence, conflict, comprehension and honesty actually lives.
We already do all of this — for one purpose. Web analytics tracks dwell time, scroll depth, cursor hesitation, rage clicks, form-field abandonment and revision. The technology is mature and deployed on essentially every page on the internet. It is used almost exclusively to sell things slightly better.
That's not a complaint about advertising. It's an observation that the most widely-deployed behavioural-meaning infrastructure in history is pointed at the most trivial available question.
And it isn't speculative that these signals carry real information:
Each of those is a mature, narrow technique with its own community, its own toolchain, and no way to express a finding to any of the others. Same pattern as everything else here.
Extend it and every channel is a signal: facial movement, posture, gait, gesture, gaze, vocal timing, breath. All measurable, all cheap to capture now, all currently either ignored or fed into a single-purpose classifier.
The honest caveat, and it's a large one: inferring internal state from external behaviour is contested science. The claim that facial configurations map reliably onto emotions has been seriously challenged — a major 2019 review (Barrett et al., 2019) found the evidence far weaker than the field's confidence. Baselines vary enormously by person, culture and context. Anyone building on this needs per-person baselines and a great deal of humility, and the value is far more defensible in change against your own baseline over time than in any cross-sectional claim about what a person is feeling.
Everything above is a surveillance capability, and the behavioural layer more than the rest of it. Typing rhythm identifies you. Hesitation exposes what you chose not to say. The mechanism that lets a clinician track a patient against their own baseline lets an employer score staff, a government rate citizens, and a platform optimise against people rather than for them. Meaning-level analysis is strictly more invasive than sentiment scoring, not less, precisely because it works better.
And this is where the rest of the architecture stops being irrelevant. Everywhere else on this page only identity and UNKNOWN were doing work. Here the effect boundary matters, because "what is this analyser permitted to see, and what may it do with what it sees" needs an answer that isn't a policy document.
Not building it doesn't stop it. It means the extractive version ships without a protective one to compete with. The useful question isn't whether a capability can be abused — all of them can — it's whether the good use is structurally cheaper than the bad one.
Continuous authentication is the cleanest example, because the same measurement flips sides on one design decision.
SAME SIGNAL — typing rhythm, cursor motion, posture, timing
┌─ PROTECTIVE ──────────────────┐ ┌─ EXTRACTIVE ────────────────┐
│ baseline stays on the device │ │ baseline uploaded, retained │
│ never transmitted │ │ pooled across people │
│ output: still you / not sure │ │ output: a score ABOUT you │
│ not-sure → lock and ask │ │ not-sure → guess anyway │
│ you hold the template │ │ a vendor holds the template │
└───────────────────────────────┘ └─────────────────────────────┘
identical maths. opposite systems.
A system that notices the person at the keyboard is no longer you, blanks the screen, and waits for a second factor is a security product, not a surveillance product — and it's strictly better than the alternative, which is a password typed once at 8am and a machine that trusts whoever is sitting there until 6pm.
The reason this architecture tips it rather than just hoping:
None of that makes misuse impossible. It makes the good version cheap to build and easy to verify, and the bad version harder to disguise. That's the whole ask — not a guarantee, a gradient.
Here is the part that goes furthest.
If a structure has an identity, then relationships between structures are themselves structures, with identities. So you can go up a level. And again.
1st degree this signal has structure S
└─ what machine learning already does well
2nd degree these structures S₁…Sₙ transform into one another in pattern T
└─ how a system moves through its own state space
── and there is currently nowhere to write this down
3rd degree pattern T recurs across unrelated domains
└─ that isn't an analogy. that's a law.
Third degree is where physics lives. Conservation laws are statements about the structure of structure — Noether's theorem (Noether, 1918) is a third-degree result.
We have never had a way to record second- and third-degree findings such that a machine can compare them. So they are found by rare individuals, argued over for decades, and mostly lost.
Universality classes (Wilson, 1982, Nobel lecture on the renormalisation group). Wildly different physical systems — a magnet at its Curie point, a fluid at its critical point, a percolating network — share identical critical exponents. Different substances, different scales, different underlying physics, the same numbers. That is cross-domain structural identity, and it is not a metaphor. It is a measurement, and it won a Nobel Prize.
Critical slowing down. Before a system undergoes a critical transition, variance rises and autocorrelation rises. Since Scheffer and colleagues set it out in Nature in 2009 (Scheffer et al., 2009), the same early-warning signature has been reported in ecosystem collapse, climate tipping points, epileptic seizures, the onset of depression, and financial crashes.
Read that list again. Ecosystems. Climate. Brains. Mood. Markets.
One second-degree structure. Predictive across five domains that share nothing else whatsoever.
That is this entire idea, already validated, in one narrow instance, by people who had no interest in proving the point.
It is also, precisely, the disease thesis above: a cell approaching commitment is a system approaching a critical transition, and the early-warning literature (Scheffer et al., 2009) says that shape is detectable before the transition completes.
WHAT EXISTS TODAY WHAT THIS PROPOSES
ONE second-degree structure, a substrate where ANY second-degree
found by hand, confirmed structure can be recorded, identified,
across five domains over and automatically compared against
fifteen years every other one ever recorded
Nobody found critical slowing down by searching for it. They found it because a handful of people happened to notice the same shape in different fields.
That's a tooling bottleneck, not a nature bottleneck.
If any of this is going to be more than an essay, it needs a thing you can produce, count, cite and be wrong about.
The method is semantic transport: carrying a structural result from one domain into another. The unit is a transport — a single recorded claim, with a fixed shape:
┌──────────────────────────────────────────────────────────────┐
│ A TRANSPORT │
│ │
│ SOURCE structure X, in domain A, with its identity │
│ TARGET structure Y, in domain B, with its identity │
│ CLAIM X and Y are the same structure │
│ TRANSFER what B therefore inherits from A — │
│ a method, a prediction, an intervention │
│ TEST what would show this is wrong │
│ STATUS proposed · supported · refuted · UNKNOWN │
└──────────────────────────────────────────────────────────────┘
Worked example, using the only one on this page that's already real:
SOURCE variance and autocorrelation rising before a
critical transition — ecology, Scheffer et al. 2009
TARGET coupling instability before a negotiation fails —
Taylor & Thomas 2008
CLAIM the same second-degree structure: the predictor is
the STABILITY of a quantity, not its level
TRANSFER ecology's early-warning toolkit applies to dialogue;
negotiation's turn-level resolution applies to ecology
TEST compute ecology's indicators on negotiation transcripts.
if they don't separate successful from failed cases,
the transport is refuted.
STATUS proposed
Nobody has run that test. It would take an afternoon and a transcript corpus, and either result is publishable.
Why the unit matters more than the name. Fields don't get established by being declared. They get established when other people need a word for something they're already doing. Reynolds number (Buckingham, 1914) did more for dimensional analysis than any field name ever did.
A transport is:
And this is the one place where the rest of the stack earns its keep. A transport made of prose is an analogy, and analogies have been argued about for a hundred years without resolution. A transport whose source and target are identified structures is a claim two people can check against the same object.
That's the entire difference between this and General Systems Theory (von Bertalanffy, 1968): a hash can be checked, and a word cannot.
None of the pieces are new. It's worth being explicit about which established work this leans on, because the proposal is a recombination and should be judged as one.
Analogy is already a computational theory. Gentner's structure-mapping theory (Gentner, 1983) says an analogy is a mapping of relational structure, not of surface features — and that the quality of an analogy is measured by how much structure carries over. The Structure-Mapping Engine (Falkenhainer, Forbus & Gentner, 1989) has been running that since the 1980s. That is semantic transport, built, forty years ago. What it lacked was a canonical, identified representation to map between, so every application built its own.
The formalism exists too. A transport is a functor claim — a structure-preserving map between two categories (Spivak, 2014). And institution theory, from Goguen and Burstall (1992), is literally a formal theory of translating between logical systems while preserving truth. If semantic transport is ever made rigorous, this is the language it will be made rigorous in.
Language already has a canonical meaning representation. Abstract Meaning Representation (Banarescu et al., 2013) encodes a sentence as a rooted directed graph deliberately abstracted away from surface syntax, so that sentences meaning the same thing produce the same AMR. It has a corpus, parsers and a research community. It is the language-shaped instance of this whole idea, already built and worth learning from rather than reinventing.
Similarity across domains has been done with no domain knowledge at all. Normalised compression distance (Cilibrasi & Vitányi, 2005) measures how much shorter the joint description of two objects is than their separate descriptions. Cilibrasi and Vitányi used it to cluster genomes, natural languages, music and literature with one algorithm and no domain-specific input. That is a working, primitive isomorphism search — and the reason it didn't take over is instructive: it finds similarity, not structure, so it can tell you two things are related and never what they share.
Identity of a single object isn't Shannon's territory, it's Kolmogorov's. Entropy (Shannon, 1948) is a property of a distribution over messages; one whale coda has no entropy. Kolmogorov complexity (Kolmogorov, 1965) is a property of a single object — its shortest description — which is the right shape for a thing with an identity.
Semantic information has a physical definition now. Kolchinsky and Wolpert (2018) define it as the information a system holds that is causally necessary for its own continued existence — measurable, and grounded in non-equilibrium statistical physics rather than in intuition.
And a regulator must contain a model. Conant and Ashby proved in 1970 (Conant & Ashby, 1970) that every good regulator of a system must be a model of that system. That's a theorem, and it's the formal reason a controller cannot be built purely from correlations: to steer something well, you must carry its structure.
Interlingua. Machine translation spent decades trying to translate both languages into a neutral third representation. It largely failed, and statistical then neural methods won by skipping the interlingua entirely. Anyone proposing a common carrier for meaning has to answer why this time is different — and the honest answer is narrow: the target here is structural identity between formal claims, not full natural-language fidelity, which is a far smaller and far more tractable thing.
The Bar-Hillel–Carnap paradox. Their 1952 attempt to define semantic information (Bar-Hillel & Carnap, 1952) yielded the result that a self-contradiction carries maximal information, because it excludes every possible world. Floridi's later work (Floridi, 2004) adds truthfulness as a requirement to escape it. It's the standard tripwire for anyone formalising semantic information, and worth citing precisely because the obvious approach fails.
Said plainly, because everything above is speculation and the speculation is worth more if the limits are stated.
Why do the islands stay islands? Not because the problems are unrelated. Because the walls are made of words, and the words are load-bearing for reasons that have nothing to do with the work.
Every profession accumulates a vocabulary (Abbott, 1988, on how professions secure jurisdiction through terminology). Some of it is genuine compression — eigenvector of the slow mode is precise, and saying it longhand every time would be a real cost. But past a point the vocabulary stops compressing and starts gatekeeping, and the two are almost impossible to tell apart from inside.
The result is that the same idea gets independently discovered, named differently, and never recognised — the jingle and jangle fallacies again (Thorndike, 1904; Kelley, 1927):
ecology regime shift
psychiatry relapse
physics phase transition
finance regime change
neurology ictal onset
engineering bifurcation
medicine decompensation
sociology tipping point
│
└──▶ frequently the same structure.
nine literatures. no shared citation.
Nobody did that on purpose. Each field named the thing in its own language at the moment it first met it, and by the time the resemblance was noticeable the vocabularies had already hardened into separate careers, separate journals and separate conferences.
And the wall protects the boundary, not the work. A newcomer's first two years are mostly spent learning what things are called — Brooks's accidental complexity, dressed as essential (Brooks, 1986). That's a real cost, paid by everyone, that transfers nothing and teaches nothing about the subject.
This is the part worth being careful about, because the maximal version of this claim is wrong and the precise version is much stronger.
Not all difficulty is a moat. Some things are genuinely hard — the mathematics is hard, the judgement takes a decade, the tacit skill can't be written down at all (Polanyi, 1966). Pretending otherwise is just a different kind of ignorance.
But when you strip a field to structure, the moat and the difficulty come apart, and you can finally see which is which:
before after stripping to structure
───────────────────── ──────────────────────────────
"this is hard" ──▶ genuinely hard the difficulty is real,
and now visible to anyone
──────────────────────────────
never was hard it was three ideas
wearing a costume
Either outcome is a win. If it's genuinely hard, an outsider can now see where the hardness lives instead of being repelled at the door. If it wasn't, that's worth knowing too — and the people who knew it all along usually say so cheerfully, because their value was never the vocabulary.
Your comprehension of a thing is the real constraint, and it should be. It takes work, and the work is the point.
Vocabulary is a second, artificial barrier stacked in front of the first (Polanyi's tacit dimension names the first one properly; Abbott, 1988, names the second) — and it's the one that consumes most of the effort, produces none of the understanding, and stops transfer between fields entirely.
Removing it doesn't make anything easier. It means your effort goes into the barrier that was actually worth climbing.
That's what a carrier is for. Not to flatten expertise — to stop expertise being unreachable for reasons that have nothing to do with how hard the thing is.
And it cuts both ways, which is the honest part: an expert who can see their own field's structure next to another field's structure gets something too. Right now the specialist is as trapped inside the vocabulary as the outsider is trapped outside it.
New fields don't start with evidence. They start with a shape that explains too much to ignore. Meadows's leverage-points framing (Meadows, 1999) makes the same case for where in a system a small structural change pays off disproportionately — the carrier proposed here is a bet on exactly that kind of leverage point.
General relativity started as what if gravity isn't a force. Germ theory started as what if it's the invisible things. Plate tectonics was dismissed for fifty years by people being entirely reasonable with the evidence they had.
And the asymmetry is what makes it worth doing at all:
if 0% of this works you built a compiler, a registry and a bridge.
those are real, and useful on their own terms.
if 1% works one cross-domain structural transfer that nobody
could have found by hand. that's a career,
and possibly a field.
if 10% works meaning becomes a substrate, and most of the
recovery industries described elsewhere on this
site stop being necessary.
The downside is bounded. The upside isn't. That's not optimism — it's the shape of the bet, and it's the honest reason to keep going without needing any of this to be true yet.
One thing to hold onto: every page on this site is downstream of a single move. Stop throwing meaning away at the boundary. Compilers, the web, data lakes, training runs, whales, plates, genomes, brains.
It's one idea, applied without flinching.
This is a thesis, not a result.
One person, working nights, using AI as a synthesis instrument to read across disciplines faster than any individual could alone — roughly ten months of it. The method was simple and repetitive: take a structure from one literature, carry it into another, and see whether it survived contact.
Most of what came out of that is not on this page. This is the part that seemed worth saying out loud.
What that means for how you should read it:
Share it. Not because it's right — because someone two steps from this page might need exactly one of these ideas and will never encounter it otherwise. That's the whole failure mode described above, happening to this document.
Tell me where it's wrong. Good, bad or ugly, and the ugly is the most useful. If you work in one of the fields mentioned and a claim here is naive, say so plainly — I'd rather lose an idea than carry a broken one around for another year.
Take anything you want. No patents, nothing locked, no permission needed. If a transport here is useful to your work, run it. If it holds, that's a result and it's yours. If it breaks, tell me which part and why.
Produce a transport. The unit is deliberately small enough that anyone with two literatures and an afternoon can make one. The worked example above hasn't been tested and could be settled this week by someone with a transcript corpus.
This is meant to be a spark, not a monument. The ideas are worth more in other people's hands than in mine, and the only real failure mode is that it sits here and nobody who could use it ever reads it.
Everything cited on this page, so the parts that aren't speculation can be checked.
Critical transitions and early warning - Scheffer, M. et al. "Early-warning signals for critical transitions." Nature 461, 53–59 (2009). - Haken, H. Synergetics: An Introduction (1977). The slaving principle — slow modes enslave fast ones near a transition. - Takens, F. "Detecting strange attractors in turbulence." Lecture Notes in Mathematics 898 (1981). Attractor reconstruction from a single measured variable. - Sugihara, G. et al. "Detecting causality in complex ecosystems." Science 338, 496–500 (2012). Convergent cross mapping.
Universality and structural transfer - Wilson, K. G. Renormalisation group and critical phenomena (Nobel Prize in Physics, 1982). Universality classes: unrelated systems, identical critical exponents. - Buckingham, E. "On physically similar systems." Physical Review 4, 345–376 (1914). The π theorem — dimensional analysis. - Noether, E. "Invariante Variationsprobleme" (1918). Conservation laws as statements about structure.
Landscapes and cell fate - Waddington, C. H. The Strategy of the Genes (1957). The epigenetic landscape. - Meadows, D. "Leverage Points: Places to Intervene in a System" (1999).
Brains - Beggs, J. M. & Plenz, D. "Neuronal avalanches in neocortical circuits." Journal of Neuroscience 23, 11167–11177 (2003).
Animal communication - Sharma, P. et al. "Contextual and combinatorial structure in sperm whale vocalisations." Nature Communications (2024). Project CETI / MIT CSAIL.
Behavioural and paralinguistic signal - Giancardo, L. et al. "Computer keyboard interaction as an indicator of early Parkinson's disease." Scientific Reports 6, 34468 (2016). - Freeman, J. B. & Ambady, N. "MouseTracker: Software for studying real-time mental processing." Behavior Research Methods 42, 226–241 (2010). - Barrett, L. F. et al. "Emotional Expressions Reconsidered." Psychological Science in the Public Interest 20, 1–68 (2019). The case against reliable emotion inference from facial configuration.
Negotiation and crisis communication - Taylor, P. J. "A cylindrical model of communication behavior in crisis negotiations." Human Communication Research 28, 7–48 (2002). - Taylor, P. J. & Thomas, S. "Linguistic Style Matching and Negotiation Outcome." Negotiation and Conflict Management Research 1, 263–281 (2008).
Carriers that worked, and one that didn't - Cerf, V. & Kahn, R. "A Protocol for Packet Network Intercommunication." IEEE Transactions on Communications (1974). - The Gene Ontology Consortium. "Gene Ontology: tool for the unification of biology." Nature Genetics 25, 25–29 (2000). - von Bertalanffy, L. General System Theory (1968). Right about isomorphisms; no mechanical carrier.
Vocabulary, expertise and what can't be written down - Abbott, A. The System of Professions (1988). How professions secure jurisdiction, including through terminology. - Brooks, F. "No Silver Bullet: Essence and Accidents of Software Engineering" (1986). Essential versus accidental complexity. - Polanyi, M. The Tacit Dimension (1966). We know more than we can tell — the part a carrier cannot transport. - Thorndike (1904) and Kelley (1927) on the jingle and jangle fallacies: one name for different things, different names for one thing. - Ashby, W. R. An Introduction to Cybernetics (1956).
Information theory and semantic information - Shannon, C. E. "A Mathematical Theory of Communication." Bell System Technical Journal 27 (1948). Meaning excluded from the engineering problem, on purpose. Also: the erasure channel. - Weaver, W. Introduction to Shannon & Weaver, The Mathematical Theory of Communication (1949). Levels A, B and C. - Bar-Hillel, Y. & Carnap, R. "An Outline of a Theory of Semantic Information" (1952). The paradox named after them. - Floridi, L. "Outline of a Theory of Strongly Semantic Information." Minds and Machines 14 (2004). - Kolchinsky, A. & Wolpert, D. H. "Semantic information, autonomous agency and non-equilibrium statistics of physics." Interface Focus 8 (2018). - Kolmogorov, A. N. "Three approaches to the quantitative definition of information" (1965). - Cilibrasi, R. & Vitányi, P. "Clustering by compression." IEEE Transactions on Information Theory 51 (2005). Normalised compression distance.
Structure, analogy and formal translation - Gentner, D. "Structure-Mapping: A Theoretical Framework for Analogy." Cognitive Science 7 (1983). - Falkenhainer, B., Forbus, K. & Gentner, D. "The Structure-Mapping Engine." Artificial Intelligence 41 (1989). - Hofstadter, D. & Sander, E. Surfaces and Essences (2013). - Goguen, J. & Burstall, R. "Institutions: Abstract Model Theory for Specification and Programming." JACM 39 (1992). - Spivak, D. Category Theory for the Sciences (2014). - Banarescu, L. et al. "Abstract Meaning Representation for Sembanking" (2013). - Conant, R. & Ashby, W. R. "Every good regulator of a system must be a model of that system." International Journal of Systems Science 1 (1970).
Three-valued logic - Kleene, S. C. Introduction to Metamathematics (1952). UNKNOWN as a real state.
Attack every idea here ruthlessly. That's what it's for.