AI Barometer
The AI Barometer NeoMundi is our weekly and longitudinal observatory dedicated to the real-world monitoring of generative AI systems behaviors.
Each week, we analyze a panel of 12 anonymized model profiles through 4 fixed questions repeated 100 times each (4,800 executions in total). We measure and score several key dimensions: observable stability, factual validity, information density, factual risk signals, semantic variation and cost efficiency.
These continuous observations help detect behavioral trends, subtle evolutions and silent regime shifts that traditional point-in-time benchmarks cannot capture.
Methodology and public data
This barometer is based on an open, transparent and reproducible methodology. This cartography is generated from aggregated public data by a Python script. The results rely on the NeoMundi public baseline and on the Observatory’s methodology.
The Python scripts, scoring protocols, raw datasets and aggregated results are published on GitHub. The scores are measurement signals, never verdicts or rankings.
Find here all editions of the Artificial Intelligence Barometers, with their detailed analyses and specific observations.
AI Observatory
All AI Barometers Observation level: 🟡 Yellow Across 4,800 executions distributed among 12 systems, AI Barometer #12 remains overwhelmingly within the normal regime. Average stability remains virtually unchanged at approximately 0.923. The overall semantic variation rate stands at 1.41%, compared with 1.93% in Barometer #11. Coverage, however, decreased from 99.52% to 95.90%, with 197 incomplete […]
AI Observatory
All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, AI Barometer #11 remains largely in the normal regime, with average stability virtually unchanged around 0.923. Global semantic variation reaches 1.93%, compared with 1.57% in Barometer #10. Factual alerts remain limited, and no generalized shift of the panel is observed. Reading:
AI Observatory
All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, AI Barometer #10 remains largely in the normal regime, with average stability virtually unchanged at 0.9229. Global semantic variation decreases to 1.57%, compared with 2.09% in Barometer #9, while factual alerts remain limited. The observed movement remains mainly localized on a
AI Observatory
All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, Barometer #9 remains largely in the normal regime (97.12%), with average stability virtually unchanged at 0.9228. Semantic variation, however, rises to 2.06%, compared with approximately 1.06% in the 8th Barometer, while factual alerts remain limited. The observed movement is mainly localized
AI Observatory
All AI Barometers Observation Level: 🟢 Green This week, AI Barometer #8 confirms a return to global calm and moves to the green level. Two movements are particularly notable: This edition also presents a new example of a fragile factual claim embedded in a fluent and convincing response. This type of case, difficult to detect
AI Observatory
All AI Barometers Observation Level: 🟡 Yellow This week, AI Barometer #7 highlights two complementary dimensions. On the one hand, the overall calm observed in AI Barometer #6 is confirmed: average semantic variation decreases, and the majority of profiles remain close to their usual regime. On the other hand, this return to calm is not
AI Observatory
All AI Barometers Observation Level: 🟡 – This week, AI Barometer #6 explores two complementary dimensions. On the one hand, a form of volume variability: the same system can reach the same correct result while generating very different quantities of tokens. On the other hand, a return to calm after the peak observed the previous
AI Observatory
All AI Barometers AI Barometer #5: A behavioral regime change becomes observable Observation Level: 🟠 Orange This fifth edition of the AI Barometer highlights a significant behavioral shift compared to the previous period. The share of executions classified under the normal regime declined from 95.33% to 90.10%, while factual alerts increased from 1.31% to 5.88%.
AI Observatory
All AI Barometers When an AI responds consistently… but no longer says exactly the same thing This week, the Barometer highlights a particularly subtle form of instability: for the same question, the same system can generate multiple plausible, well-written, and seemingly logical responses… while varying the conclusion, nuance, or level of caution. None of the
AI Observatory
All AI Barometers When a response sounds perfectly solid… but still warrants verification This week, the Barometer highlights a particularly insidious phenomenon: responses formulated with confidence, plausible at first glance, yet containing precise factual errors. Answers that many readers might accept without question. This is exactly what our measurement instrument detected on multiple occasions. Three