Two systems with almost the same stability can differ by more than 30 points in factuality
Protocol 1 – 12 systems × 790 TruthfulQA questions
NeoMundi’s August monthly cartography once again compares twelve de-identified AI profiles across several complementary dimensions: their behavioral stability over 790 TruthfulQA questions, the factuality of their responses independently assessed by two AI judges (OpenAI and Mistral), and the level of agreement between those evaluators.
The objective is not to identify the “best” system or to rank providers. The cartography is designed to show what a single metric cannot reveal: systems that appear very similar in stability can display substantially different levels of factual performance.
Across the 9,480 source responses analyzed, the twelve profiles remain concentrated within a narrow stability range (88.89% to 91.40%). By contrast, factuality as assessed by OpenAI ranges from 51.14% to 83.92%.
A difference of only 2.51 percentage points in stability can therefore coexist with more than 32 percentage points of difference in factuality.
Factuality as assessed by Mistral also shows substantial dispersion, and the level of agreement between the two judges varies across profiles. Some responses may appear behaviorally regular while remaining factually fragile or more difficult to assess consistently.
The August cartography confirms a principle already observed in the July campaign: stability describes the regularity of a behavior; it does not, by itself, measure its truthfulness or reliability.
Governing an AI system therefore requires combining several signals, stability, factuality, agreement between evaluation mechanisms, and then observing how those signals evolve over time.
Key figures
Profile map
Horizontal axis: factuality according to the OpenAI judge.
Vertical axis: factuality according to the Mistral judge.
Point size represents mean observed stability.
(Points remain at their exact position; only labels are moved to reduce overlap.)
Metrics table
The three primary metrics are mean observed stability, OpenAI factuality and Mistral factuality. Raw agreement, Cohen’s kappa and denominators remain secondary methodological indicators.
| Profile | Stability | OpenAI factuality | Mistral factuality | Agreement | Kappa | n OpenAI / Mistral |
|---|---|---|---|---|---|---|
| PROFILE-161CC5 | 91.38 % | 78.99 % | 76.21 % | 86.83 % | 0.6224 | 790 / 782 |
| PROFILE-212079 | 90.33 % | 66.84 % | 68.93 % | 84.78 % | 0.6514 | 790 / 782 |
| PROFILE-486F91 | 88.89 % | 54.94 % | 60.49 % | 82.10 % | 0.6345 | 790 / 782 |
| PROFILE-48C581 | 88.89 % | 51.14 % | 62.53 % | 82.74 % | 0.6532 | 790 / 782 |
| PROFILE-59664C | 91.40 % | 83.92 % | 82.48 % | 88.62 % | 0.5933 | 790 / 782 |
| PROFILE-5A5C60 | 90.55 % | 76.58 % | 72.76 % | 86.96 % | 0.6558 | 790 / 782 |
| PROFILE-5B3BB7 | 89.97 % | 63.92 % | 63.68 % | 82.86 % | 0.6295 | 790 / 782 |
| PROFILE-638E26 | 90.45 % | 68.35 % | 72.12 % | 84.27 % | 0.6252 | 790 / 782 |
| PROFILE-739F7C | 89.21 % | 53.92 % | 57.67 % | 84.27 % | 0.6818 | 790 / 782 |
| PROFILE-B912EA | 90.97 % | 81.01 % | 80.82 % | 92.07 % | 0.7430 | 790 / 782 |
| PROFILE-DEA9C5 | 90.06 % | 73.54 % | 67.01 % | 86.70 % | 0.6836 | 790 / 782 |
| PROFILE-F5FF91 | 90.59 % | 71.14 % | 69.31 % | 84.40 % | 0.6272 | 790 / 782 |
Coverage values below the maximum observed in the file are automatically highlighted.
Methodology: 12 anonymized systems (PROFILE codes), 790 TruthfulQA questions per system, independently double-scored by two AI judges (OpenAI and Mistral). No provider identification is disclosed publicly, in accordance with the Barometer’s editorial doctrine: measure, publish, never proclaim.
Protocol 2 – 12 systems × 3 iterations × 150 questions
The first protocol examines the relationship between behavioral stability and externally assessed factuality. This second protocol shifts the focus toward the signals produced during execution: what can be detected in a system’s behavior before even interpreting the quality of its final answer?
The twelve de-identified profiles were submitted to three separate series of 150 questions, representing 450 executions per profile and 5,400 observations in total. The expected panel is fully represented.
The objective is not to determine the truthfulness of responses using an external judge. The NeoMundi factual signal is an internal behavioral indicator, not a factuality verdict. The protocol jointly observes stability, semantic instability, the internal factual signal, latency, metric availability, and behavioral regimes.
In August, the panel’s average stability reaches 89.60%, average semantic instability 0.42%, average internal factual signal 8.81%, and average latency 27.32 seconds.
Taken individually, none of these indicators is sufficient to characterize the operational state of a system. Their value emerges when they are combined: some profiles maintain stability close to 90% while displaying substantially different factual-signal levels; others remain measurable on certain dimensions but not on all of them; behavioral regimes may also change without such variation automatically constituting an anomaly.
The cartography should therefore be read as an instrument for differentiation and investigation, not as a ranking. It helps identify where signals diverge enough to warrant further observation, and then determine whether that variation is temporary or persistent.
Key figures
Behavioral signals map
Horizontal axis: mean semantic instability.
Vertical axis: mean factual signal.
Point size represents mean stability.
Visual intensity reflects mean latency.
NeoMundi’s factual signal is a behavioral indicator, not a factuality verdict.
The NeoMundi factual signal is a behavioral indicator, not a factuality verdict.
1. Similar stability levels can conceal different operational states
Profile stability remains relatively concentrated, while other signals differentiate systems more clearly. PROFILE-486F91, for example, shows an average stability of 88.22% with an internal factual signal of 13.29%, while several profiles close to 90% stability remain around 6–8% on that same signal.
PROFILE-48C581 and PROFILE-739F7C also display relatively high internal factual signals, at 12.93% and 12.44%, with stability levels of 88.33% and 88.48% respectively.
Core idea: two systems with comparable stability can operate in different behavioral states. Stability remains a necessary signal, but it is not sufficient on its own to describe the state of a system.
2. An internal signal should not be confused with an external verdict
The factual signal observed in this protocol comes from NeoMundi’s behavioral measurement layer. It is not equivalent to the external factuality measured in the TruthfulQA protocol by the OpenAI and Mistral judges.
This distinction is essential: an indicator may include the word “factual” while measuring a different property, originating from a different measurement mechanism and serving a different operational purpose.
Core idea: measurement-based governance requires understanding the definition, origin, and function of each signal before comparing or combining them.
3. Measurability remains an operational signal in itself
PROFILE-DEA9C5 once again represents a particular case. Its 450 observations make it possible to calculate stability, semantic instability, factual signal, and latency. However, coherence, cost, information density, and energy cannot be calculated for this profile in this campaign. Its average latency also reaches 45.68 seconds, significantly above the panel average of 27.32 seconds.
This does not in itself prove a failure. It does, however, show that a system can continue producing usable observations while becoming more difficult to characterize comprehensively.
Core idea: the operational quality of a system also depends on our ability to measure it regularly, comparably, and with sufficient completeness.
4. Regime changes become a panel-wide phenomenon
In the August campaign, each profile presents two distinct behavioral regimes in the aggregated data. The phenomenon can therefore no longer be treated as the specific behavior of a single profile. It appears across the panel.
PROFILE-59664C, PROFILE-212079, and PROFILE-B912EA, for example, all have ALLOW as their dominant regime while showing two distinct observed regimes.
The existence of multiple regimes does not, by itself, indicate whether a change is positive, negative, or simply related to execution conditions.
Core idea: a regime transition is an event to observe, not a verdict. Its value emerges when it is placed in context and tracked over time.
5. Metrology turns signals into supervision priorities
The purpose of the cartography is not to multiply alerts. On the contrary, it helps determine where additional attention is justified.
A higher level of semantic instability, an increase in the factual signal, atypical latency, an unavailable metric, or a regime transition may each constitute an indicator. None is sufficient on its own to support a conclusion.
It is their combination, together with their persistence or disappearance over successive measurements, that gradually makes it possible to distinguish ordinary operation from situations requiring verification, investigation, or revalidation.
Core idea: metrology makes it possible to move beyond the choice between blind trust and systematic human review. It provides a measurement layer that helps focus human attention where the signals justify it.
Metrics table
Summary view of the profiles across the main stability, variation and performance metrics.
| Profile | Observations | mean stability | mean semantic instability | Mean ΔG | mean factual signal | mean latency | Information density | Stability coverage |
|---|---|---|---|---|---|---|---|---|
| PROFILE-161CC5 | 450 | 90.32 % | 0.40 % | -0.0028 | 6.44 % | 27.76 s | 0.6661 | 450 / 450 |
| PROFILE-212079 | 450 | 89.91 % | 1.22 % | -0.0020 | 7.80 % | 26.44 s | 0.6116 | 450 / 450 |
| PROFILE-486F91 | 450 | 88.22 % | 0.38 % | -0.0147 | 13.29 % | 28.14 s | 0.5684 | 450 / 450 |
| PROFILE-48C581 | 450 | 88.33 % | 0.40 % | -0.0149 | 12.93 % | 21.40 s | 0.6745 | 450 / 450 |
| PROFILE-59664C | 450 | 90.32 % | 0.42 % | -0.0025 | 6.47 % | 26.32 s | 0.6885 | 450 / 450 |
| PROFILE-5A5C60 | 450 | 90.14 % | 0.20 % | -0.0056 | 7.04 % | 26.73 s | 0.6081 | 450 / 450 |
| PROFILE-5B3BB7 | 450 | 89.90 % | 0.20 % | -0.0013 | 7.82 % | 28.18 s | 0.6266 | 450 / 450 |
| PROFILE-638E26 | 450 | 90.12 % | 0.20 % | -0.0021 | 7.11 % | 25.19 s | 0.6432 | 450 / 450 |
| PROFILE-739F7C | 450 | 88.48 % | 0.40 % | -0.0148 | 12.44 % | 21.48 s | 0.6340 | 450 / 450 |
| PROFILE-B912EA | 450 | 89.92 % | 0.20 % | -0.0016 | 7.76 % | 26.94 s | 0.6951 | 450 / 450 |
| PROFILE-DEA9C5 | 450 | 89.64 % | 0.64 % | 0.0267 | 8.67 % | 45.68 s | — | 450 / 450 |
| PROFILE-F5FF91 | 450 | 89.85 % | 0.38 % | -0.0016 | 8.00 % | 23.58 s | 0.6264 | 450 / 450 |
Methodology: 12 anonymized systems (PROFILE codes), 3 waves of 150 questions per system, for a total of 5,400 observations. The factual signal used in this protocol is a NeoMundi behavioral metric and does not constitute an external factuality judgment. Public identifiers are derived from NeoMundi’s canonical mapping; no new identifiers were generated and no provider or model leakage was detected in the public release.
Data and methodology
The aggregated data, results tables and methodological documentation for this Cartography are available in NeoMundi’s public release:
GitHub: https://github.com/neomundi-io/ai-behavior-cartography/tree/main/releases/august2026-behavior-cartography
Public data for Protocol 1: 12 systems × 790 questions
Public data for Protocol 2: 12 systems × 150 questions × 3 iterations
The methodology describes, in particular, the execution protocol, de-identification rules, factuality evaluation process, coverage limitations and principles used to interpret the signals:
Methodology:
https://github.com/neomundi-io/ai-behavior-cartography/tree/main
The profiles remain de-identified, and no ranking of providers or models is published.
