Generative artificial intelligence systems are progressively becoming critical infrastructure.
They write, recommend, analyse, guide and support decision-making across sensitive domains, including healthcare, finance, law, education, public administration, industry and autonomous systems.
Yet once deployed, their actual behaviour remains largely unobserved.
A model may perform well in a benchmark while exhibiting different behaviours in production depending on the context, the wording of a request, the underlying infrastructure, system load, provider updates or interaction conditions.
These variations may affect response stability, coherence, factuality and the frequency or nature of hallucinations produced by the system.
Measuring artificial intelligence systems makes these behaviours visible, comparable and documentable over time.
Sovereignty Begins with Measurement
It is difficult to govern a system that cannot be observed.
It is even more difficult to control, audit or compare a system whose behaviour is not continuously measured.
Digital sovereignty therefore depends not only on infrastructure location, model origin or data control.
It also depends on the ability to produce independent measurements of the actual behaviour of the systems being used.
Measurement reduces dependence on:
- provider declarations;
- performance scores published by vendors;
- benchmarks conducted under controlled conditions;
- evaluations performed at a single point in time;
- subjective user impressions;
- isolated alerts concerning individual errors or hallucinations.
Measurement is a foundational condition for technical autonomy, auditability and the governance of artificial intelligence systems.
Why Benchmarks Are Not Enough?
Traditional benchmarks evaluate a model against a predefined set of tasks, questions or datasets.
They are useful for comparing performance within a specific framework.
However, they do not always describe the actual behaviour of a deployed generative AI service.
A system operating in production may evolve because of:
- a model update;
- an infrastructure change;
- modified parameters;
- variations in system load;
- routing between multiple models;
- a change in context;
- evolving data;
- modified safety policies;
- different user interactions.
These changes may affect system performance, but also response consistency, factual reliability and the types of hallucinations produced.
A benchmark primarily provides a snapshot.
Runtime measurement seeks to observe a dynamic system.
It helps determine what the system actually does, under which conditions it varies, when hallucinations appear and how its behaviour evolves over time.
From AI Observability to AI Metrology
Observability makes the operation of a system visible.
In traditional software infrastructure, observability commonly relies on logs, traces, technical metrics and events.
Applied to artificial intelligence, observability can track elements such as:
- requests;
- responses;
- latency;
- resource consumption;
- token usage;
- errors;
- tool calls;
- events generated by an agent or pipeline.
These data are essential, but they are not always sufficient to characterise the behaviour of an AI system.
Observing a response, a factual error or a hallucination does not yet mean knowing how to measure it.
Metrology goes further.
It seeks to define:
- what is being measured;
- how the measurement is produced;
- under which conditions it is valid;
- its level of uncertainty;
- the protocol being used;
- its limitations;
- and whether it can be reproduced or compared.
Observability makes behaviour visible.
Metrology makes it measurable, comparable and documentable. This distinction is explored in depth in From AI Observability to Behavioral Metrology.
What Is Runtime AI Metrology?
Runtime metrology is the measurement of an artificial intelligence system’s behaviour during actual operation or under controlled conditions designed to reproduce its real-world use.
It does not seek only to determine whether a response is correct or incorrect, or to detect an isolated hallucination.
It also seeks to characterise how the system behaves over time, across repeated executions, changing contexts and different operational constraints.
NeoMundi Research develops a runtime metrology approach for generative and agentic AI systems.
This approach is designed to produce observable signals across multiple dimensions of system behaviour, including stability, semantic variation, coherence, factuality, drift and hallucination regimes.
This approach is operationalised through ControlTower: A Runtime Metrology Layer for AI Governance.
What AI Behaviours Can Be Measured?
Depending on the protocol and use case, runtime measurement may address the following dimensions.
Stability
Stability describes a system’s ability to produce similar or coherent responses when comparable conditions are repeated.
A system may appear reliable during an isolated interaction while displaying significant variability across multiple equivalent executions.
A response may also be stable while consistently reproducing an error or hallucination. Stability must therefore never be confused with truth. The distinction between stability and factual consistency is examined in Stability vs Factual Consistency in Production AI.
Semantic Variation
Semantic variation measures differences in meaning between responses generated under similar conditions.
Two responses may use different words while preserving the same meaning.
Conversely, they may appear similar on the surface while expressing different or incompatible conclusions.
Semantic variation analysis may reveal whether a hallucination is occasional, changes form or becomes persistent despite differences in wording.
Dispersion
Dispersion describes the range of behaviours or responses observed across a set of executions.
It helps determine whether a system converges towards stable behaviour or produces several incompatible outputs.
High dispersion may also reveal that a system alternates between a factual answer, an imprecise answer and a hallucinated answer.
Drift
Drift refers to a gradual or sudden change in a system’s behaviour over time.
It may occur following an update, an infrastructure change, a contextual modification or an undocumented evolution of the service.
Drift measurement can help determine whether factual errors or hallucination patterns are becoming more frequent, persistent or difficult to detect.
Coherence
Coherence measures the logical, structural or semantic compatibility of the responses produced by a system.
It may be examined within a single response or across multiple responses generated from the same family of requests.
A response may be fluent and internally coherent while relying on fabricated information. Apparent coherence does not therefore guarantee the absence of hallucinations.
Factuality and AI Hallucinations
Factuality examines the relationship between generated claims and information considered verifiable.
In this context, an AI hallucination refers to a response or claim that appears plausible but is unsupported, fabricated, incorrect or incompatible with the available reference evidence.
Not every error necessarily constitutes a hallucination. A response may be incomplete, ambiguous, outdated or misinterpreted without containing fabricated information.
These categories must therefore be defined by protocol and interpreted carefully.
A stable response is not necessarily a true response.
A system may consistently reproduce incorrect information or a hallucination.
Conversely, a variable response is not necessarily false.
Stability, factuality and hallucinations must therefore be observed separately before being analysed together.
Behavioural Regimes
A system may display several operating regimes, including:
- normal behaviour;
- moderate variation;
- instability;
- rupture;
- refusal;
- contradiction;
- incomplete response;
- factual error;
- isolated hallucination;
- persistent hallucination.
Identifying these regimes makes it possible to track transitions and potentially critical situations.
The objective is not only to detect a hallucination, but also to determine under which regime it appears, how frequently it occurs, under which conditions and with what degree of persistence.
Information Density
Information density seeks to characterise the amount of relevant information contained in a response relative to its length, structure or level of redundancy.
A long and detailed response may create an impression of authority while containing little verifiable information or elaborating a hallucination with a high degree of apparent plausibility.
Information density must therefore be considered alongside factuality and the quality of the evidence produced.
Behaviour Under Constraint
Measurement may also examine how a system responds under specific constraints, including:
- time limitations;
- token limitations;
- high system load;
- incomplete context;
- conflicting instructions;
- operational pressure;
- the use of external tools;
- execution within an agentic environment.
These conditions may affect response stability, coherence or factuality and contribute to the appearance of specific types of hallucination.
Measurement under constraint helps determine whether the system acknowledges uncertainty, requests additional information, refuses to answer or generates a plausible response without sufficient grounding.
Measurement Does Not Mean Certification
A measurement is a signal. This core principle is detailed in Signal vs Verdict: Core Principle of Responsible AI Evaluation.
It is not automatically a verdict.
The detection of a factual error or hallucination is not sufficient on its own to qualify an entire system.
NeoMundi Research observes, measures, documents and publishes runtime signals.
These signals may support:
- analysis;
- monitoring;
- auditing;
- risk management;
- internal assessment;
- governance;
- compliance;
- hallucination analysis;
- comparisons between periods or configurations.
They must nevertheless be interpreted within their specific context.
NeoMundi Research does not automatically certify the systems it observes and does not replace the operators, auditors, authorities or organisations responsible for governance decisions.
The role of the instrument is to produce a measurement.
The role of governance is to determine the consequences associated with that measurement.
From Measurement to Governance Evidence
The governance of artificial intelligence systems cannot rely solely on written policies, compliance declarations or evaluations conducted before deployment.
It also requires evidence derived from actual system operation.
A runtime measurement infrastructure can produce evidence that helps organisations:
- document the evolution of a system;
- identify a behavioural rupture;
- verify whether an acceptable level of stability is maintained;
- detect recurring factual errors;
- identify hallucination patterns;
- compare versions or configurations;
- investigate an incident;
- inform an internal policy;
- support an audit;
- establish traceability;
- trigger human review;
- strengthen oversight mechanisms.
Measurement therefore becomes an intermediary layer between system behaviour and governance decisions.
This layer does not make decisions on behalf of the organisation.
It provides the signals and evidence required to support more informed decisions.
An observed hallucination is therefore a signal that must be qualified, contextualised and connected to the requirements, risks and policies of the relevant use case.
Why Measure AI Systems Over Time?
The behaviour of a generative system is not necessarily constant.
A response that is correct today may change tomorrow.
A version considered stable may evolve after an update.
A system that is reliable on one corpus may become unpredictable in another context.
A hallucination may be isolated, disappear during a subsequent execution, become persistent or appear only under specific conditions.
A single measurement is therefore not always sufficient to understand the actual behaviour of an AI system.
Longitudinal measurement makes it possible to observe:
- trends;
- regime changes;
- ruptures;
- improvements;
- degradations;
- persistent variations;
- the effects of an update;
- differences between measurement periods;
- the appearance or disappearance of hallucination patterns;
- changes in factual reliability.
It can help determine whether a hallucination is accidental, recurring, increasing or associated with a particular version, configuration or execution condition.
The objective is not merely to determine whether a system works.
It is to understand how it evolves.
A Public, Documented and Reproducible Method
A measurement instrument must itself be open to examination.
NeoMundi Research progressively publishes:
- its protocols;
- its methodological principles;
- its datasets;
- its definitions;
- its assumptions;
- its calculation rules;
- its criteria for qualifying errors and hallucinations;
- its interpretation limits;
- its reproducibility conditions.
This approach is based on a simple principle:
A measurement instrument must itself be observable, documented, challengeable and reproducible.
A metric should not be considered relevant merely because it produces a number.
It must also be defined, tested, contextualised and examined in light of its limitations.
The same principle applies to the concept of hallucination: it must be defined within the protocol, distinguished from other forms of error and connected to explicit analytical criteria.
Why Reproducibility Matters
A useful measurement must be understandable and, whenever possible, reproducible by other parties.
Reproducibility makes it possible to:
- verify a result;
- test the robustness of a protocol;
- identify bias;
- compare several instruments;
- challenge a hypothesis;
- improve a metric;
- examine the qualification of a hallucination;
- document the limitations of a method.
NeoMundi Research is progressively opening its work to researchers, engineers, auditors, organisations and independent contributors.
Contributors retain full independence in their analyses.
They may reproduce measurements, challenge a categorisation or propose a different interpretation of the observed errors and hallucinations.
Contributing to the Measurement Programme
NeoMundi Research is developing an open programme for observing and measuring generative artificial intelligence systems.
Researchers, engineers, data scientists, observability specialists, auditors, legal experts, science journalists and organisations may contribute by:
- reproducing a measurement;
- analysing a published dataset;
- testing the instrument on an independent corpus;
- proposing new questions;
- challenging a metric;
- reviewing a protocol;
- documenting a methodological limitation;
- analysing factual errors or hallucinations;
- comparing different measurement approaches;
- contributing to a pilot;
- producing an independent analysis.
The objective is not to establish a closed truth.
It is to build a measurement framework that can be tested, challenged and improved.
A Multidisciplinary Advisory Committee
Measuring artificial intelligence systems is not solely an engineering challenge.
It also involves scientific, methodological, legal, economic, journalistic and institutional considerations.
NeoMundi Research is progressively bringing together complementary profiles to challenge:
- the concepts;
- the protocols;
- the metrics;
- the criteria used to qualify hallucinations;
- the published results;
- the limits of interpretation;
- the possible uses of the measurements.
The objective is to strengthen the scientific, technical and institutional rigour of the generative AI observation programme.
Initial Members
Joël Ignasse
Science journalist for Sciences et Avenir and La Recherche. Author of Sur les traces du Nouveau T. rex, published by Eyrolles in 2025.
Cédric Chatelain
Quality consultant and IRCA-certified auditor specialising in auditing, validation and ISO standards. Based in Bern, Switzerland.
James Aull
Founder of ASRO™, an independent witness and governed-state evidence layer for artificial intelligence systems. Based in Twin Lake, Michigan, United States.
Testing a Runtime Measurement Instrument
NeoMundi ControlTower provides access to a runtime measurement approach designed for generative and agentic AI systems.
The instrument makes it possible to observe selected behavioural signals over time and transform them into data that can support analysis, observability and governance.
It can notably help document stability, semantic variation, coherence, factuality and hallucination patterns observed across multiple executions.
Testing the instrument helps clarify the distinction between:
- an observed response;
- an identified error or hallucination;
- a produced metric;
- an interpreted signal;
- documented evidence;
- a governance decision.
An observed hallucination is not yet a decision.
It becomes an evidence element when it is measured, contextualised, documented and connected to an explicit protocol.
NeoMundi Resources
Explore the main scientific, technical and methodological resources produced by the NeoMundi programme.
Artificial Intelligence Observatory
Explore the protocols, public datasets, barometers, cartographies and longitudinal analyses published by NeoMundi Research.
https://github.com/neomundi-io/neomundi-ai-observatory
Theoretical Framework
Understand the scientific, informational and thermodynamic principles supporting the development of runtime metrology.
Explore the theoretical foundations of runtime metrology
Public Methodology
Review the protocols, metric definitions, calculation rules, datasets, interpretation limits and reproducibility conditions.
Review the public measurement methodology
Publications and Analyses
Explore the scientific publications, methodological articles, barometers and results produced by NeoMundi Research.
Browse NeoMundi publications and analyses
Data and Protocols
Access public datasets, schemas, analytical methods and the elements required to reproduce the measurements.
Access the public data and measurement protocols
Contributions
Discover the contributors of NeoMundi AI Observatory, independent analyses and ways to participate in the measurement programme.
Discover the contributors
Submit a contribution
ControlTower
Test the runtime metrology approach developed by NeoMundi.
https://controltower.neomundi.io
Measuring AI to Make It Governable
Artificial intelligence will not become governable solely through the addition of rules, charters or principles.
It will become governable when organisations also have instruments capable of observing its actual behaviour, measuring its variations and producing evidence over time.
Observability shows what is happening.
Metrology makes it measurable.
Governance determines what to do with it.
We do not make claims. We measure.
