Methodology
Spring 2026 UK Edition
This note is written for institutional readers who require methodological transparency. A shorter public overview is available on the main Research Intelligence page.
1. Methodological Position
1.1 What this methodology refuses to do
Moirai does not reduce research identity to a single number. Institutional research is too heterogeneous — in field mix, scale, and how it changes over time — for any single scalar to honestly describe it. The methodology that follows is built to refuse that collapse: structure, position, evolution, and evidence are surfaced as separable dimensions, comparable across editions, and never folded into a single ranking verdict.
1.2 What Moirai Research Intelligence provides
Moirai Research Intelligence gives universities a structural, comparable, and defensible view of their research identity — enabling leaders to explain where their institution stands, what it is becoming, and how it can be understood beyond rankings or self-reported claims.
This is not a ranking. It is a structural reading of an institution's research output, computed from public bibliographic sources and surfaced through a transparent, versioned methodology.
1.3 Position relative to existing systems
Moirai is not affiliated with REF, THE, QS, ARWU, Leiden Ranking, or Nature Index. It complements these sources; it does not replace them.
1.4 The two axes of the methodology
The methodology is organised along two complementary axes.
Structure — the static topology of an institution's research at a given Edition: communities, fields, archetype, position. Structure answers what is the shape of this institution's research right now? Structure is described in §3 through §8.
Evolution — the temporal continuity of research observed without rewriting historical Editions: the daily activity of research, the year-by-year shifts within a published Edition, the movement across published Editions, and the persistence of research-community identity across Editions. Evolution answers how does this institution's research live and change over time? Evolution is described in §9.
The two axes are complementary, not hierarchical. A reader interested in what an institution is should read Structure surfaces; a reader interested in how the institution is changing should read Evolution surfaces. Neither axis is collapsed to a single trajectory verdict, in keeping with §1.1.
2. The Edition Cycle
Moirai publishes in annual Spring Editions with one Autumn Refresh mid-cycle.
| Phase | Months | What happens |
|---|---|---|
| Spring Edition release | May–July | New 5-year window data, fully recomputed indicators, public release |
| Operating window | June–October | Edition is the published reference for the cycle |
| Autumn Refresh | October–November | Citation data updated to capture late accruals; edition identity does not change |
| Cycle planning | December–March | Next edition's parameters and source-layer freezes |
Why annual. Research structure does not meaningfully change quarter-to-quarter. Annual cycles match how universities plan, how research offices report, and how funding bodies operate.
Why Spring. May–July aligns with UK academic decision windows: end-of-academic-year board reviews, pre-summer strategic planning, and June graduation-cycle institutional reflection. Earlier cycles risk incomplete current-year data; later cycles miss the decision window entirely.
Each Spring Edition is permanently archived as a snapshot — Editions are forever. Once an Edition is published, no field of any indicator computed for that Edition is ever modified. The Autumn Refresh updates citation accruals as an associated refresh layer under the same Edition family, not as a mutation of the Edition's published values. Daily incremental ingestion, retroactive abstract enrichment, and historical backfills feed the live infrastructure described in §9 but never mutate the snapshot a customer was shown at Edition publication.
This is the technical property that makes evolution honest. Without it, "year-over-year change" would be self-referentially retroactive: today's reading of last year's number could differ from last year's published reading of last year's number, and no comparison could be defended. Edition immutability is the prerequisite for the Evolution axis (see §9.2).
Historical Editions remain queryable. Cross-Edition movement becomes first computable at the Spring 2027 publication (see §9.3 and §14.8).
3. The Spring 2026 UK Edition
3.1 Window and scope
- Data window: 2021–2025 (5 years, rolling)
- Geographic scope: United Kingdom
- Universe: 243 UK institutions selected by the inclusion rule below
- Comparison universe for percentile-based metrics: this 243-institution UK universe
Two scope contexts. This Edition uses a rolling 5-year window (2021–2025). Moirai also maintains a fixed long-window baseline view (2014–2025) used for cross-institution comparability; because the two views use different denominator windows, the same coverage measure can differ between an institution's Edition page and its baseline page. This is by design, not an inconsistency.
3.2 Inclusion rule
A UK institution is included in the Spring 2026 UK Edition universe if it satisfies:
country = "United Kingdom"in OpenAlex- AND (
has_mri = 1in Moirai's reference table OR explicitly featured for the cycle) - AND it is not classified as an NHS Trust, Hospital Foundation, or other clinical-care institution
NHS Trust and Hospital Foundation exclusion. 18 NHS Trusts and Hospital Foundation entities are excluded from the Spring 2026 UK Edition universe. These institutions have substantial publication output in OpenAlex but their publication patterns are dominated by clinical workflow rather than research-strategic choice. Including them would distort percentile-based comparison metrics for academic research-intensive universities, which are the primary subjects of the Edition. NHS-affiliated research conducted within partner universities (for example, Imperial College London or King's College London) is captured through those academic institutions in the normal way.
The selection list is frozen for the cycle and recorded in institution_list.json (sha256-versioned).
3.3 Eligibility for ISP scoring
Within the 243-institution universe, ISP composite scores are computed only for institutions with output_5y ≥ 2,000 (5-year publication count of 2,000 or more).
For Spring 2026 UK:
- 138 institutions are ISP-eligible (output_5y ≥ 2,000)
- 105 institutions are ISP-ineligible (output_5y < 2,000) and remain in the universe with structural attributes (Archetype, community structure, GFP) but no ISP composite score
Smaller institutions are not omitted; they are reported with the indicators their data volume supports. The ISP composite eligibility threshold exists because the Item Response Theory model underlying the Quality component requires sufficient citation observations to produce stable estimates.
4. Indicator Definitions
Every metric on every page references these definitions. No synonyms, no ad-hoc rewording.
4.1 Public language
| Public name | What it means | Range |
|---|---|---|
| Research Potential | Composite indicator of institutional research strength potential | 0.0–1.0 |
| Research Influence | How an institution's outputs are received in their field-year context | Percentile (0–100) |
| Research Capacity | 5-year research output, expressed as percentile | Percentile (0–100) |
| Top-end Output | Share of an institution's output landing in the global top 10% of its field-year | Percentile (0–100) |
| Research Diversity | Effective number of research communities | Percentile (0–100) |
| Structural Resilience | Diversity remaining after the three largest communities are removed | Percentile (0–100) |
| Research Coverage | Share of 5-year output represented in the analytical graph | Percent |
4.2 Public name vs internal computation
"Research Potential" is Moirai's public expression of Institutional Research Potential (ISP): a structural indicator of an institution's capacity to sustain and amplify high-quality research. It is not a ranking of past output. It is shown as an institutional structural indicator, not an ordinal league-table rank, and it is not a substitute for REF, THE, QS, or other rankings.
4.3 Field Baseline Signal Profile
The Field Baseline Signal Profile is a structural product surface introduced in v1.0.3 of this methodology. It complements the ISP composite by providing a per-field and per-community view of an institution's research output relative to a global field-year reference distribution.
4.3.1 Two-layer signal model
Moirai's institutional Position surface uses two complementary layers:
Field-layer signals tell you where the institution sits within its disciplinary peer group: against 138 UK research universities, and against 2,240 globally tracked institutions. The signal is the rate at which the institution's papers reach the top-decile citation threshold for their field and year.
Community-layer signals tell you about the institution's research communities — clusters of researchers and collaborators identified through co-authorship network structure. Each community's signal is the rate at which its in-field papers reach the global top-decile citation threshold for the community's primary field.
These two layers honour different questions. Field-layer answers "where does this institution sit among peers?" Community-layer answers "how does this specific community of researchers compare to the global field baseline?" Both questions are surfaced honestly; neither replaces the other.
Note on counts: field-layer is institution-wide (every paper attributed to the institution in the 5-year window contributes); community-layer is graph-attributed (a paper contributes only if it belongs to a community in the co-authorship graph). The institution-wide count is the larger of the two and is what readers see in the field-layer view; the graph-attributed subset is what drives the community-layer bands. The relationship between the two is the basis for the Coverage strip (§4.3.4 below).
4.3.2 26-field taxonomy
Moirai v1.x uses a 26-field taxonomy drawn from OpenAlex Topics at the Field level. This taxonomy is the stable v1.x field set for the Spring 2026 launch. Subfield and topic-level additions in OpenAlex Topics do not affect Moirai's v1.x signals because v1.x reads field-level labels only. Future field-level changes in OpenAlex Topics would trigger field-mapping review and are tracked separately.
The 26 fields are: Agricultural and Biological Sciences; Arts and Humanities; Biochemistry, Genetics and Molecular Biology; Business, Management and Accounting; Chemical Engineering; Chemistry; Computer Science; Decision Sciences; Dentistry; Earth and Planetary Sciences; Economics, Econometrics and Finance; Energy; Engineering; Environmental Science; Health Professions; Immunology and Microbiology; Materials Science; Mathematics; Medicine; Neuroscience; Nursing; Pharmacology, Toxicology and Pharmaceutics; Physics and Astronomy; Psychology; Social Sciences; Veterinary.
These 26 labels are identical to the Scopus subject-area top-level naming because OpenAlex Topics is built on Scopus subject codes. Moirai reads the labels from OpenAlex Topics; Scopus is the labelling lineage, not Moirai's data source.
4.3.3 Top-tail signal — what it measures
For each community, Moirai computes the share of the community's in-field papers (over the 5-year window 2021–2025) that reach or exceed the global field-year top-decile citation baseline. We call this the top-tail signal because it measures the upper tail of research impact, not the average. A community with a top-tail signal of 31% has 31 papers above the global field-year p90 baseline for every 100 in-field papers it produced in the 5-year window.
We deliberately avoid the phrase "high-impact" because "impact" in the UK research context refers primarily to societal impact (REF Impact case studies), which is a different and richer concept. Our top-tail signal is a citation-based measure of upper-tail research density.
Per-paper threshold. Each paper's "above" status is tested against the field-year p90 baseline for the paper's own primary field and year — not against the community's primary field. The community's top-tail signal restricts to its in-field papers: those whose own primary field matches the community's primary field. This restriction ensures the community's signal reflects the upper-tail density of its core disciplinary output, not a heterogeneous mix.
Three-tuple public display. Community rates are always shown as a three-tuple rate · count · denominator. "31% — 38 of 124" is unambiguous: 38 in-field papers above the global Social Sciences field-year p90, out of 124 in-field papers in the 5-year window. Rate alone, band alone, or percentile alone are forbidden on public surfaces because they lack a quantitative anchor.
Six community-layer bands. Each community receives one of:
| Band | Public label |
|---|---|
strong_top_tail_signal | Strong top-tail signal |
above_expected_top_tail | Above expected top-tail signal |
near_expected_top_tail | Near expected top-tail signal |
below_expected_top_tail | Below expected top-tail signal |
limited_top_tail_signal | Limited top-tail signal |
insufficient_sample | Insufficient sample |
The first five community-layer bands are derived from calibrated top-tail-rate thresholds for the Spring 2026 UK distribution. The sixth (insufficient_sample) applies when the community has fewer than three in-field papers in the 5-year window or when the community's primary field is unmappable (§14.7 below).
Five field-layer bands. Each (institution, field) pair receives a parallel UK-universe band and Global-universe band, drawn from a separate threshold set:
| Band | Public label |
|---|---|
upper_quartile | Upper quartile |
above_median | Above median |
around_median | Around median |
below_median | Below median |
lower_quartile | Lower quartile |
UK-universe thresholds are computed across 138 ISP-eligible UK institutions; Global-universe thresholds are computed across the 2,240 institutions globally tracked with sufficient output. Both threshold sets are frozen for the edition.
4.3.4 Abstract-enriched subset
Counts and rates in Moirai's institutional surfaces are based on the abstract-enriched subset of OpenAlex papers — papers that have been processed through Moirai's research extraction pipeline. As abstract coverage rises edition-over-edition, the metric scope widens proportionally. The Coverage strip on each institution surface reports the current coverage in plain numbers (graph-attributed papers out of total 5-year output, year-by-year).
Users comparing Moirai counts to OpenAlex / Scopus / Web of Science direct exports may see lower numbers; this is expected and reflects the abstract-enrichment scope, not an under-counting.
Abstract enrichment is the gating filter for inclusion in field-layer and community-layer signals. A paper without an abstract in Moirai's analytical store is not counted in either layer. This is the metric scope contract: every count, rate, and band is defined over the abstract-enriched subset.
4.3.5 Sample confidence
For each community, Moirai labels the sample size with one of four levels: Robust sample (30 or more in-field papers in 5 years), Moderate sample (10–29), Limited sample (3–9), or Insufficient sample (fewer than 3, no rate reported). These thresholds are calibrated against the Spring 2026 UK distribution and may adjust in future editions. The labels are stable; the numeric cut-points evolve.
We do not currently report statistical confidence intervals. If distribution analysis shows that narrow-sample bands dominate positive signals, we will add Wilson confidence intervals as a launch requirement in a future edition.
The internal calibration enum is sample_confidence with four values: robust_sample / moderate_sample / limited_sample / insufficient_sample. The labels are frozen; the numeric cut-points are calibration-governed and may move while the four-label scheme remains stable.
4.3.6 What the surfaces do not show
Institutional Position surfaces in v1 do not surface:
- Individual researcher names, ORCID, or anchor authors
- Paper titles, abstracts, or paper-level evidence drawers
- Citation-momentum trajectories (deferred to cross-Edition Evolution surfaces beginning Spring 2027 — see §9.3)
- Single-number institutional verdicts or league-table positions
These exclusions are deliberate. Strategic Access engagements unlock researcher-level and paper-level evidence in conversation, not in the public surface, and only after editorial review by Moirai analysts.
Strategic Access teaser. Where a community surface includes a Strategic Access reservation, the teaser drawer header is "Strategic Access module" and the locked teaser body is the frozen single-line wording:
Anchor researchers and representative evidence are reserved for Strategic Access engagements.
The teaser body carries no numeric counts, no quantity language (few, several, many, multiple), and no implicit cardinality (a team of researchers, a body of papers). The frozen wording is published verbatim.
Schema reservation. Strategic Access fields (anchor authors, paper drilldown) are reserved in the output schema with status schema_reserved_not_computed and emission rule emit_when_empty: never. The public-tier API response does not emit these keys (not even as empty arrays). Tier visibility is enforced server-side; CSS/JS hide patterns are not acceptable as substitutes.
5. ISP v4 — Composite Formula
ISP combines three signals:
- Capacity (C): How much research the institution publishes. Measured as percentile of 5-year research output within the edition's universe.
- Quality (Q) — publicly surfaced in the Intelligence Portal as Research Influence: How research outputs are received in their field-year context. Measured as percentile of latent ability θ from an Item Response Theory model over field-year citation thresholds.
- Yield (Y): What share of an institution's research lands in the global top 10% of its field-year. Stabilised against publication-volume noise via credibility shrinkage.
The composite formula:
ISP = C^0.60 × RQP^0.40
RQP = Q^0.70 × Y^0.30
Capacity has the largest weight (0.60). Quality and Yield together — the Research Quality Potential (RQP) factor — supply the remaining 0.40.
5.1 Yield (Y) construction
Plain-English summary. Yield measures whether an institution produces more top-decile research than would be expected for its field mix and publication volume. Institutions with a small but exceptionally cited research portfolio receive credit for genuine top-end output; smaller institutions are protected against random variation by being partially anchored to the global expected rate.
Technical detail.
yield_adj = λ × top_tail + (1 − λ) × 0.10
λ = output_5y / (output_5y + 20,000)
Y = Percentile(yield_adj, edition_universe)
The credibility shrinkage parameter λ weights an institution's observed top-tail share (papers above the field-year P90 threshold) against the global expected pass rate of 0.10. Smaller institutions receive heavier shrinkage toward the global expectation; larger institutions are weighted toward their own observed share.
The expected pass rate is exactly 0.10 because the field-year P90 threshold is calibrated so that 10% of papers in each eligible field-year bucket exceed it by construction.
5.2 Cascade lock and reproducibility
ISP v4 is computed in four stages, with each stage's output cryptographically locking the next:
| Stage | Output | Cascade input |
|---|---|---|
| f6 — Field-year reference | Percentile thresholds per field-year bucket | source ES |
| g1 — IRT ability | Per-institution latent quality θ + ability_norm | f6 sha256 |
| g2 — Top-tail | Per-institution top_tail (papers above P90 / quality_papers) | g1 sha256 |
| g3 — ISP composite | Per-institution Capacity, Quality, Yield, RQP, ISP | g2 sha256 |
Every stage records its script_sha256 and input_version (the upstream sha256). Re-running with the same inputs produces byte-identical outputs. Cascade integrity is verified at runtime; mismatch fails the run before any write.
6. Methodology Disclosures
6.1 RQP is field-year normalised quality density
RQP (the Quality × Yield factor) is a field-year normalised quality-density signal. It is not a measure of total research volume, institutional scale, or overall university strength.
ISP combines RQP with scale-adjusted research capacity (the Capacity component, weight 0.60) precisely so that focused high-density signals are not interpreted as total institutional strength. An institution with a small but exceptionally cited research portfolio will have high RQP; a large institution with broad mainstream output will have high Capacity. ISP balances these.
6.2 IRT model field-aggregation (Spring 2026 UK)
The Item Response Theory model in g1 produces a single latent ability θ per institution, aggregated across fields rather than per-field. This methodology choice has known consequences:
- Institutions with heterogeneous field mixes may have quality signals smoothed across disciplines, rather than separately estimated within each field. This can make the Quality component less sensitive to field-specific peaks and troughs than a per-field θ model would be.
- Large comprehensive universities may show their disciplinary breadth more clearly through the Capacity component than through the single aggregated Quality component, because Capacity captures output volume across all fields directly while Quality compresses field-specific variation into a single estimate.
This is a methodological boundary, not a defect. Spring 2027 cycle will introduce IRT v3 with field-specific θ to address this. Until then, the ISP composite weighting (Capacity 0.60) ensures that scale and breadth remain primary drivers for large multi-disciplinary institutions.
6.3 ISP composite — confidence intervals
ISP confidence intervals are not emitted in the Spring 2026 UK Edition. While the underlying IRT model produces credible intervals on the latent ability θ, these do not propagate cleanly through the non-linear composite formula C^0.60 × (Q^0.70 × Y^0.30)^0.40. Bootstrap-based composite CIs are deferred to a future cycle.
6.4 2025 citation data are provisional
Citation counts for 2025 publications are still accruing at the Spring 2026 publication date. Moirai uses OpenAlex snapshot 2025-11-12, with incremental updates through 2026-03-30, as its citation source for this edition.
The Autumn 2026 Refresh (October–November) will update citation counts against a later OpenAlex snapshot. The edition identity (publish: 2021_2025_spring_uk) does not change; only the citation freshness improves. The methodology version, score version, archetype version, and underlying community structure remain frozen for the cycle.
6.5 Topic-Level analysis (strategic briefs)
A subset of Spring 2026 UK Edition deliverables — specifically institutional strategic briefs — uses Topic-Level analysis in addition to the Phase 1 portal indicators. Topic-Level analysis and Phase 1 portal indicators are computed independently and reported as complementary lenses.
How Topic-Level analysis is computed. Topic-Level analysis uses the OpenAlex Topic taxonomy. For each Topic-Year cohort, the global 90th-percentile citation threshold is calculated, and institutional papers are evaluated against the corresponding Topic-Year threshold. An above-P90 rate of 1.0× represents the global expected rate. A rate of 2.0× indicates an institution produces twice the global expected share of top-decile papers in that Topic-Year. Threshold calculations are year-matched to avoid cross-year citation-age bias; recent-year citation windows are naturally less mature and are treated accordingly.
How this differs from Phase 1 portal indicators. Phase 1 portal indicators (ISP, Capacity, Quality, Yield, Archetype) use field-level aggregation across 26 OpenAlex-derived field-level categories with an ordered-response Item Response Theory model. The two lenses operate at different levels of granularity: field-level for portal indicators, topic-level (a finer taxonomy nested within fields) for strategic briefs.
Status and forward path. Topic-Level analysis is currently used as a forward-looking diagnostic layer for strategic briefs and is under evaluation for integration into the core methodology in a future release cycle. Until that integration, Topic-Level findings are reported separately from the ISP-anchored Phase 1 indicators and should be read as a complementary lens rather than a replacement for the Phase 1 percentile-anchored measures. Both lenses are reproducible and use the same underlying OpenAlex source snapshot for the Edition.
6.6 Field Baseline Signal disclosure
The Field Baseline Signal Profile (see §4.3) reports community-level signals relative to a field-year reference baseline frozen for the cycle. Three methodological notes apply to public-tier readers:
1. Reference baseline is frozen per edition. The field-year p90 thresholds used to compute top-tail signals come from the f6_field_year_reference (sha256 prefix 344b29d5, frozen for Spring 2026 UK Edition; 130 buckets across the 26 fields × 5 years 2021–2025). The same reference is used for ISP Yield (Y) computation (see §5.1). Re-running the field baseline build with the same inputs against the same frozen reference produces byte-identical outputs.
2. Community structure is shared across the cycle. As noted in §8.1, the underlying community map (Leiden communities computed once per major cycle) is shared across editions within the cycle. The Field Baseline Signal Profile inherits this structural baseline: what changes per edition is the band each community receives, not the community membership itself. This is methodologically intentional — it allows edition-over-edition comparison of community signal movements while keeping the structural skeleton stable. Citation-momentum trajectory (community signal movement edition-over-edition) becomes first computable on cross-Edition Evolution surfaces from the Spring 2027 publication (see §9.3 and §14.8).
3. Data inclusion contract is fixed and visible. Every field-baseline build records its inclusion contract in per-institution audit metadata. For v1 of the Field Baseline Signal Profile, the contract requires: abstract-enriched papers only; institution attribution via the OpenAlex authorship link; OpenAlex logid present (an OpenAlex-anchored record). Crossref-ingested records pending OpenAlex institution resolution are not included in the v1 field-layer build; their inclusion is part of writer-level v1.0.1, which will extend the Crossref ingestion writer with ROR → OpenAlex institution resolution and historical backfill (see §14.7). The year window is 2021–2025. The contract is identical across all 138 ISP-eligible UK institutions in this edition; auditors verify by reading the metadata of any institution's output, not by reading source.
These three disclosures are the contract under which Position-tab signals can be cross-checked across institutions and across editions.
7. Archetype Classification
Each ISP-eligible institution is classified into one of four research-structure archetypes based on community concentration (Gini index) and effective community count (effective N):
| Archetype | Gini | Effective N |
|---|---|---|
| Broad & Balanced | Low | High |
| Established & Diverse | High | High |
| Specialist & Cohesive | High | Low |
| Focused & Emerging | Low | Low |
Spring 2026 UK Edition uses the archetype version uk_v1_med_log_2021_2025 — UK-baseline median split on the log-transformed effective N axis.
Archetype is a categorical classifier, not a quality judgment. All four archetypes describe distinct strategic profiles. A Specialist & Cohesive institution is not "below" a Broad & Balanced one; they reflect different research strategies and serve different mandates.
8. The Structural Layer
8.1 Research Communities
Moirai identifies research communities by applying Leiden community detection to a per-institution graph built from semantic vector similarities of research findings extracted from publication abstracts.
- Vectorisation: SciNCL embeddings (768-dimensional, scientific-text-trained)
- Vector store: Qdrant (cosine similarity, 50-nearest-neighbour graph)
- Community detection: Leiden algorithm (resolution-tuned for institutional scale)
- Structural baseline: 2020–2024 community detection run, frozen for cycle
Community structure is treated as a slow-moving structural baseline. It is computed once per major cycle and shared across editions within the cycle. Institution-level annual indicators (ISP, Archetype, percentile rank, Atlas position) are recomputed each Spring Edition; the underlying community map is not. The community baseline is intentionally lagged by one cycle to ensure structural stability and avoid rebuilding community maps from citation-fresh, still-maturing current-year data.
8.2 Coverage definitions
Three structural measures, in increasing strictness:
| Measure | Definition |
|---|---|
Analysis Coverage (scope_pct) | Share of 5-year output that entered the analytical graph |
Structured Coverage (structured_pct) | Share of 5-year output assigned to a Leiden community |
Display Share (display_pct) | Share of the analytical graph represented in V3-tier published communities. This is a display share of what is shown at the high-level structural layer, not a measure of total research coverage. |
Lower values at any level reflect either (a) genuine output diversity beyond the analytical scope, (b) abstract availability gaps in the source data, or (c) communities below the V3 threshold for public display. Coverage is reported transparently on every institution page.
8.3 Global Field Positioning (GFP)
GFP shows where an institution's research distribution differs from the global field distribution:
- Strengths: Fields where the institution's share of output exceeds the global share, weighted by field citation impact
- Strategic Gaps: Fields where the global share substantially exceeds the institution's share
Clinical and multidisciplinary fields (Medicine, Nursing, Dentistry, Veterinary, Pharmacology, Health Professions, Multidisciplinary, Unknown) are excluded from GFP comparison. These are excluded because their publication patterns are dominated by clinical workflow rather than research-structural choice; including them would mask genuine research-strategic differentiation.
GFP is a shared structural-layer indicator. It uses the same community structure as the rest of the edition's structural layer; the comparison reference is the global field distribution from global_field_stats.json.
9. The Evolution Layer
Sections 3 through 8 describe the Structural Layer — the static topology of an institution's research at a single Edition. This section introduces the Evolution Layer — how that topology changes over time. The Evolution Layer is constituted by four layers, each answering a distinct question, each with a distinct data lineage, and each with a distinct surface availability in the Spring 2026 UK Edition.
9.1 The four layers
L1 — Activity Layer. L1 surfaces what is happening in an institution's research as of today. It draws on the daily Crossref and OpenAlex ingestion streams to identify clusters of newly published or recently active research, recent-activity heatmaps by field and by community, and a daily-pulse summary of where the institution's research is moving in the short term. L1 reads change every day; there is no Edition snapshot inside L1. The customer surface for L1 launches mid-2026; the underlying infrastructure is already in production.
L2 — Structural Drift. L2 looks inside a single published Edition and asks how the institution's research profile shifted across the years within that Edition's window. For the Spring 2026 UK Edition (window 2021–2025), L2 will surface year-by-year field-mix shifts, community-size trajectories, and ISP-component (C / Q / Y) decomposition across the five years. L2 is derived strictly from the frozen Edition snapshot and is bit-identically reproducible from the published cache. L2 is the first Evolution surface and is in implementation.
L3 — Edition Continuity. L3 compares two or more published Editions and measures how an institution's position has moved across them — percentile movement, archetype migration, field-position shifts. L3 is derived strictly from frozen Edition snapshots and inherits Edition Immutability from §2: yesterday's reading of last Edition's number equals today's reading of last Edition's number, however much time has passed. L3 is not yet computable in the Spring 2026 UK Edition because only one Edition is published; L3 becomes first computable at the Spring 2027 publication (see §9.3).
L4 — Identity Persistence. L4 is the deepest layer of the Evolution architecture. Across two or more published Editions, Moirai matches research communities (the structural primitives introduced in §8.1) to track which communities persist, which emerge, which dissolve, and which merge or divide. The output is a small set of identity events — persist, birth, death, merge, split — reported as discrete community-level observations, not as a single identity-continuity index. L4 is the layer that lets an institution see whether the research groupings that defined it in one Edition are still the research groupings that define it in the next. L4 becomes first computable at the Spring 2027 publication (see §9.3).
9.2 Edition Immutability — the prerequisite
Section 2 commits the Edition Cycle to immutability: once published, the values computed for an Edition are never changed. That commitment is what makes the Evolution Layer honest. Without it, the Layer would not measure change — it would measure the difference between two readings of the same point. Daily incremental ingestion, retroactive abstract enrichment, and any other live-data update feeds L1 (the live activity layer) but never mutates a published Edition snapshot. L2, L3, and L4 derive from frozen Editions; their outputs are reproducible bit-identical from the same Edition cache. This is also what permits an external methodology reviewer to verify any Evolution-layer claim months or years after it was first surfaced: the Edition snapshot the claim was computed against is still byte-identical to the Edition snapshot the reviewer is reading.
9.3 The Spring 2027 anchor
L3 and L4 are not features awaiting development. They are computations awaiting a second endpoint. A trajectory across Editions cannot be computed when only one Edition is published. The Spring 2027 publication — the first Edition where a paired second snapshot exists — is when L3 and L4 become first computable. The Trajectory tab in the Spring 2026 UK Edition Intelligence Portal surfaces L3 and L4 as locked capability previews; these previews are not deferred features. They are the only honest surface available before the second Edition is published.
9.4 Surface availability in the Spring 2026 UK Edition
| Layer | What it answers | Surface status in Spring 2026 UK Edition |
|---|---|---|
| L1 — Activity | What is happening in research today | Infrastructure live; customer-facing surface launching mid-2026 |
| L2 — Structural Drift | How the institution's profile shifts across years within this Edition | First computable Evolution surface; in implementation |
| L3 — Edition Continuity | How position changes across Editions | Becomes first computable at Spring 2027 publication; locked capability preview in Portal Trajectory tab |
| L4 — Identity Persistence | Whether research communities persist, emerge, dissolve, or transform across Editions | Becomes first computable at Spring 2027 publication; locked capability preview in Portal Trajectory tab |
Methodological anchor: Moirai Evolution Layer v0.1, frozen 2026-05-21, sha8 9830b8bf.
10. Comparison Universes
| Universe | Inclusion | Used for |
|---|---|---|
| Spring 2026 UK Edition (243) | UK institutions per §3.2 | Default for percentile metrics on UK pages |
| ISP-eligible UK (138) | UK universe with output_5y ≥ 2,000 | Default for ISP composite + Capacity / Quality / Yield percentiles |
| Global MRI | All institutions worldwide with has_mri = 1 | Available as a secondary comparison context |
Every universe used on a page is named explicitly. Institutions are never compared across universes without disclosure.
11. Data Sources and Lineage
| Source | Role | Use in Spring 2026 UK Edition |
|---|---|---|
| OpenAlex | Primary bibliographic, citation, institution, topic and field source | Primary calibration source for ISP, Topic-Level briefs and structural indicators (snapshot 2025-11-12 + incremental through 2026-03-30) |
| Crossref | DOI metadata enrichment and daily coordination layer | Supplementary metadata layer; not a primary ISP calibration source |
| ORCID | Researcher identity reference | Not used in ISP v4.0 scoring |
| UKRI Gateway to Research | UK funding evidence | Contextual evidence layer; not used in ISP v4.0 scoring |
Edition lineage version pins (frozen):
- Methodology:
ris_v1.1(this document) - Score:
v4.0(unchanged from v1.0) - Archetype:
uk_v1_med_log_2021_2025 - Baseline:
baseline_v2_uk_2021_2025_lag1_202604 - Findings extraction:
v6 - Claims extraction:
v6.1 - Evidence layer:
v1.0.1 - Graph build:
v7 - Community detection:
leiden_v7_2020_2024 - Field-year reference:
344b29d5(f6 sha256 prefix) - Field Baseline Signal Profile spec (base v0.4.2):
d69cf7ec(sha256 prefix) - Field Baseline Signal Profile spec (v0.4.3 overlay — OpenAlex Topics direct field assignment):
2b842588(sha256 prefix) - Field Baseline Signal Profile spec (v0.4.4 overlay — v1 OpenAlex-anchored inclusion contract):
531bde08(sha256 prefix) - Field Baseline Signal Profile methodology: v1.1 (this document)
- PI Anchor 5 patch (pair-update, v0.4.3):
f42e5a32(sha256 prefix) - Moirai Evolution Layer spec (v0.1, public projection anchor for §9):
9830b8bf(sha256 prefix) - Tier visibility config:
tier_visibility_config_v1.json(server-side enforcement; locked-tier fields never traverse the network)
12. Governance Disclosures
12.1 Face validity
Upper-tail and lower-tail ISP outputs were reviewed through a documented methodology review process before production sign-off. Large research-intensive institutions cluster in the top quartile by Capacity; ISP composite ranges differ by archetype as expected by the methodology design.
12.2 Component coherence
ISP component pairwise correlations are computed and reported in the edition's metadata (see _metadata.component_correlations in the published artifact). For Spring 2026 UK:
- Maximum correlation between any single component (C, Q, Y) and the composite ISP did not exceed the 0.95 single-factor-domination warning threshold.
- All correlations are reported transparently rather than collapsed into a single derived score.
12.3 Determinism
g3 is fully deterministic. Same inputs produce byte-identical outputs. The script sha256 is recorded with every published artifact for reproducibility audit.
12.4 What we will NOT claim
Moirai does not claim:
- That ISP correlates with REF, THE, QS, or any other ranking
- That ISP predicts grant success, recruitment outcomes, or institutional strategy
- That community structure represents intent (it represents observed publication patterns)
- That field-year P90 thresholds reflect "good" or "important" research universally
- That higher ISP is "better" without context — different archetypes, mandates, and disciplinary mixes serve different ends
Moirai claims only that ISP, Archetype, and community structure are reproducible structural readings of an institution's published research output, computed transparently from public sources.
13. Edition Lifecycle Status
Spring 2026 UK Edition status (as of 2026-05-15 launch).
13.1 Edition identifiers (three-layer naming)
| Layer | Identifier | Role |
|---|---|---|
| Edition (data) | 2021_2025_v1 | Parent baseline shared across regional Editions in the same cycle |
| Edition (publish) | 2021_2025_spring_uk | Canonical published Edition for this UK release |
| Edition (alias) | 2021_2025_v1_uk | User-facing alias on the portal |
The data layer captures the underlying field-reference baseline. The publish layer captures the regional release identity. The alias layer is the human-readable label exposed on portal URLs and Edition View Bars.
13.2 Cascade lineage
Each computational stage's sha256 cryptographically locks the next, providing end-to-end reproducibility:
| Stage | Status | sha256 |
|---|---|---|
| Universe | validated (243 institutions) | recorded in institution_list.json |
| f6 — field-year reference | validated | 344b29d5… |
| g1 — IRT ability | validated | c874ceca… |
| g2 — top-tail | validated | 87e39695… |
| g3 — ISP composite | validated | cb6588b1… |
| projection_sync — published to institutional strength projection table | completed (2026-05-02) | run_id recorded in pipeline_run_state |
| GFP edition artifacts | materialised (240 institutions) | 3 institutions recorded as expected no-GFP cases |
The sha256 values shown above are validated as of methodology sign-off and are re-confirmed against the published per-edition artifact metadata at public launch. The authoritative values for any edition are recorded in that edition's _metadata block in the published JSON artifact.
13.3 Edition lifecycle states
planning → computing → validated → active. Spring 2026 UK Edition entered validated between 2026-05-01 (g3 composite sign-off) and 2026-05-02 (projection_sync stage completion); transitions to active on public launch 2026-05-15.
14. Known Limitations and Forward Work
14.1 IRT field-aggregation (Spring 2026)
See §6.2. Per-field θ is planned for a future methodology cycle, with Spring 2027 as the current target.
14.2 ISP confidence intervals
See §6.3. Bootstrap composite CIs scheduled for a future cycle.
14.3 Citation freshness for 2025 publications
See §6.4. Autumn 2026 Refresh resolves.
14.4 Topic-Level integration into core methodology
See §6.5. Topic-Level analysis is currently a strategic-brief diagnostic layer; integration into core Phase 1 portal indicators is under evaluation for a future release cycle.
14.5 Data-source coverage
OpenAlex provides broad coverage of peer-reviewed journal literature across most fields, but coverage varies by output type, language, discipline, and source availability. Coverage is weaker for:
- Books and book chapters in Arts & Humanities
- Conference papers in Computer Science (workshop-tier proceedings)
- Non-English research outputs in some social science disciplines
The Spring 2026 UK Edition is primarily calibrated around verified academic publication outputs. Non-traditional outputs such as patents, industry research, clinical trial registries, and some grey literature are not yet used as primary calibration signals for this edition.
14.6 Cross-sector coverage
Moirai is currently focused on institutional research identity based on verified academic outputs. Over time, we are also building toward a broader cross-sector view of knowledge development, so universities can better understand their place not only within academia but within wider research ecosystems. No release timeline is committed.
14.7 Field Baseline Signal — unknown_field subsplit, inclusion contract, coverage
This subsection documents four disclosures specific to the Field Baseline Signal Profile (§4.3).
unknown_field row — first-class entry with three-way subsplit. The Field Baseline Signal Profile preserves an unknown_field row as a first-class entry rather than discarding output that cannot be mapped to one of the 26 canonical fields. The row is subsplit into three categories:
| Subsplit | Cause | Display label |
|---|---|---|
outside_field_taxonomy | Paper has an OpenAlex Topics field assignment, but the assigned value lies outside the 26 canonical field labels used in this edition. Rare under current OpenAlex Topics; defends against future taxonomy version differences. | Outside Field Taxonomy |
primary_topic_missing | Paper is recognised by OpenAlex but its primary topic has not been assigned. Includes pre-2023 papers from before the OpenAlex Topics taxonomy rollout, certain document types not topic-classified (some preprints, datasets), and a small residual awaiting Topics enrichment. | Awaiting Topics enrichment |
oas_lag | Paper is OpenAlex-anchored (logid present, attribution timestamp recorded) but its OpenAlex Topics classification has not yet propagated through to the field assignment used by the build. Distinct from primary_topic_missing in that the paper was recently anchored to OpenAlex and Topics enrichment runs on a slower cadence than core record landing. Resolves typically within one OpenAlex Topics enrichment cycle. | Topics propagation pending |
The three subsplits reflect distinct causes with distinct stability profiles. outside_field_taxonomy is rare and structural. primary_topic_missing is the largest practical bucket — partly stable (older papers OpenAlex will likely not retrofit) and partly transient (newer papers awaiting enrichment). oas_lag is transient — it resolves over the OpenAlex Topics enrichment cadence.
A high primary_topic_missing share is informational, not a defect — it typically reflects an institution's heavy historical legacy output that pre-dates OpenAlex Topics. A high oas_lag share is a temporary state, and the next edition will reduce this number.
Inclusion contract for v1. The v1 Field Baseline Signal Profile includes published output that reached Moirai through two OpenAlex-anchored ingestion paths:
- OpenAlex-primary modern. OpenAlex
logidpresent, DOI present, OpenAlex attribution timestamp set. Full field anchoring; counted in both field-layer and community-layer signals when in-field. - OpenAlex-primary pre-DOI.
logidpresent, DOI absent. Field anchoring via OpenAlex Topics; counted in both layers.
Crossref-ingested records pending OpenAlex institution resolution — published output that reached Moirai through the daily Crossref incremental layer (DOI present, logid absent) without yet being resolved to OpenAlex institution identifiers — are not included in the v1 field-layer build. These records carry institution attribution as free-form affiliation strings rather than as resolved OpenAlex institution identifiers, which prevents the field-layer build from attributing them to a specific institution under the v1 inclusion contract.
The forward path is writer-level v1.0.1: the Crossref ingestion writer will be extended with ROR → OpenAlex institution resolution at the point of ingestion, and a historical backfill will retro-fit institution identifiers to existing Crossref-ingested records. Once writer-level v1.0.1 lands, these records become structurally eligible for inclusion in subsequent Field Baseline Signal Profile builds. The methodology version does not change as a consequence of this writer-level change; what changes is the size of the included evidence pool, transparently recorded in per-institution audit metadata at each subsequent build.
Build-time data hygiene. Each build excludes a small set of historical placeholder records that carry a literal sentinel string in the OpenAlex identifier field (a data artefact from an early Crossref incremental writer). The exclusion is mechanical and does not affect institutional counts.
Coverage edition-over-edition. Coverage in the field baseline view increases edition-over-edition as abstract enrichment continues and as OpenAlex attribution catches up to recent Crossref-ingested output. The public-facing note on this surface is:
Coverage rises edition-over-edition as abstract enrichment continues.
This is a deliberate editorial choice: rather than report a single static coverage number, the surface acknowledges that the analytical graph is a living dataset and that the institution's representation in it strengthens with each cycle.
14.8 Cross-Edition trajectory
The Evolution Layer's L3 (Edition Continuity) and L4 (Identity Persistence) sub-layers, introduced in §9.1, become first computable at the Spring 2027 publication. Before a second Edition exists, cross-Edition trajectory cannot be computed under any methodology. The Spring 2026 UK Edition Intelligence Portal surfaces L3 and L4 as locked capability previews in the Trajectory tab; these previews are the honest surface available before the Spring 2027 second endpoint exists.
At the Spring 2027 publication, the first cross-Edition diff (L3) and the first community-matching pass for identity events (L4) will be computed against the Spring 2026 UK Edition as the first endpoint. The methodology version may advance at that point if L3 / L4 outputs introduce indicators not present in v1.x; otherwise, the Edition publication carries forward the v1.x methodology with the addition of the cross-Edition observations.
15. Methodology Versioning and Updates
This methodology is versioned. Material changes to formulas, thresholds, eligibility rules, or universe definitions trigger a new methodology version (ris_v1.1, ris_v2.0, etc.) and are disclosed both in this document and in the per-institution edition metadata.
Substantive expansions of the methodological surface that do not change formulas, thresholds, eligibility, or universe — such as the introduction of the Evolution Layer in v1.1 — also trigger a methodology version increment so that external readers can identify the expansion by version number alone.
Editorial corrections (typo fixes, clarifications without semantic change) are tracked in the page revision history at the bottom of this document.
16. Transparency and Use
The composite formula, shrinkage parameters, and methodology disclosures in this note are published for methodological transparency and institutional review. Academic citation, methodological discussion, and non-commercial verification are welcomed.
The underlying data products, derived indicators, edition lineage, software infrastructure, institutional pages, and research intelligence outputs are produced by Moirai Research Intelligence. Commercial use, redistribution, or product integration requires prior written permission from Moirai.
17. Contact
Questions about methodology, edition data, or institutional inclusion:
For institutional licensing, custom briefings, or peer-reviewed methodology discussion:
- [email protected] (subject line: "Methodology")
Revision History
| Version | Date | Notes |
|---|---|---|
| v1.0 | 2026-05-15 | Initial public methodology note for Spring 2026 UK Edition |
| v1.0.3 | 2026-05-15 | Field Baseline Signal Profile surface added (§4.3, §6.6, §13.7). 26-field taxonomy from OpenAlex Topics at the Field level. Two-layer signal model (field-layer institution-wide; community-layer graph-attributed). 6-band community top-tail signal + 5-band field-layer bands (UK / Global universes). 4-level sample confidence. Three-way unknown_field subsplit (outside_field_taxonomy / primary_topic_missing / oas_lag). Track B / DOI-anchored inclusion policy. Strategic Access teaser frozen wording. Pair-update with PI Anchor 5 v0.4.3 patch (f42e5a32). Spec authority: d69cf7ec (base) + 2b842588 (v0.4.3 source-path amendment overlay). Score version unchanged. Edition data unchanged. v1.0.1 and v1.0.2 versions skipped (never deployed). |
| v1.1 | 2026-05-21 | Methodology surface expansion adding the Evolution axis (L1–L4). Page-1 opening rewritten as methodological-trust commitment ("Moirai does not reduce research identity to a single number"). §1 Mission restructured into four sub-sections (§1.1–§1.4). §2 Edition Cycle final paragraph strengthened with explicit Edition Immutability framing ("Editions are forever"). §6.6 third bullet rewritten for OpenAlex-anchored inclusion contract honesty (writer-level v1.0.1 trajectory). Single Q → Research Influence bridge gloss added at first §5 Q introduction. §4.1 Public language table harmonised with deployed Portal labels and Governance v1.0.3 §3.7 canonical mapping (Research Quality → Research Influence; Research Scale → Research Capacity); definitions and percentile ranges unchanged. §4.3.6 third bullet + §6.6 second note: citation-momentum trajectory deferral wording updated for v1.1 narrative consistency (former "deferred to v1.1+ when two editions exist" replaced with anchor to cross-Edition Evolution surfaces beginning Spring 2027, cross-referenced to §9.3 and §14.8; §6.6 also drops the inaccurate "Spring 2026 + Autumn 2026 Refresh" pairing per MEL v0.1 §3.1 framing of Autumn Refresh as associated refresh layer, not new Edition). New §9 The Evolution Layer (four sub-sections, surface-availability table, public projection of MEL v0.1 sha8 9830b8bf). §14.7 (renumbered from §13.7) updated inline: oas_lag definition updated for v1 OpenAlex-natural meaning; "Track B / DOI-anchored inclusion" subsection replaced with v1 OpenAlex-only honest inclusion contract and writer-level v1.0.1 forward path; Track B vocabulary swept from public methodology surface in favour of neutral "Crossref-ingested records pending OpenAlex institution resolution". New §14.8 Cross-Edition trajectory / Spring 2027 anchor. §11 Data Sources and Lineage version block synchronised (Methodology line bumped to ris_v1.1; FBSP spec v0.4.4 overlay sha8 531bde08 added; MEL v0.1 anchor sha8 9830b8bf added). §9–§16 renumber cascade applied (renumbered to §10–§17). Pair-anchors: MEL v0.1 (9830b8bf); Two-Axis Product Architecture (CG R0 2026-05-21); Edition Immutability (PI Strong Anchor 2026-05-21); Anti-Dimensional-Collapse Principle (PI Strong Anchor 2026-05-20). Score version ISP v4.0, archetype version uk_v1_med_log_2021_2025, formula C^0.60 × (Q^0.70 × Y^0.30)^0.40, eligibility threshold output_5y >= 2,000, universe definitions (243 UK / 138 ISP-eligible) all unchanged. Spec authority chain: d69cf7ec (base v0.4.2) + 2b842588 (overlay v0.4.3) + 531bde08 (overlay v0.4.4). Supersedes deferred v1.0.4 lineage (6720161d content absorbed; 4c0d6724 design-superseded by Gate 1 memo 58b531f9 and Gate 2 patch 23d93dd0). |