AI Observatory

The NeoMundi AI Observatory is our public research program dedicated to the continuous, transparent and reproducible evaluation of generative AI systems.

We regularly publish thermodynamic cartographies and in-depth observations, including the NeoMundi AI Barometer, which tracks de-identified model profiles over time across multiple dimensions: stability, factual validity, informational density, factual-risk signals, semantic variation and cost efficiency. These publications reveal behavioral trends and silent regime changes that traditional one-shot benchmarks cannot capture.

All our work is built on open methodologies and publicly available data. Generation scripts, scoring protocols and raw datasets are released on GitHub, allowing researchers, developers and the wider AI community to audit, reproduce and build upon our observations.

Students at Champlain College Saint-Lambert discovering NeoMundi’s AI Weather during an awareness session.
AI Observatory

NeoMundi Invited to Present AI Weather to Students at Champlain College Saint-Lambert

The NeoMundi team is pleased to take part in an educational initiative designed to help students better understand AI Weather and the reliability issues associated with generative artificial intelligence systems. This initiative is exclusively educational. Its purpose is to help students become better-informed users of artificial intelligence: to understand that model behavior can vary, learn […]

AI Barometer #12 – Yellow level – Coverage down, global stability maintained – NeoMundi AI Observatory
AI Observatory

AI Barometer 12, August 31 to September 6, 2026

All AI Barometers Observation level: 🟡 Yellow Across 4,800 executions distributed among 12 systems, AI Barometer #12 remains overwhelmingly within the normal regime. Average stability remains virtually unchanged at approximately 0.923. The overall semantic variation rate stands at 1.41%, compared with 1.93% in Barometer #11. Coverage, however, decreased from 99.52% to 95.90%, with 197 incomplete

AI Barometer #11 – Green level held, global stability – Localized semantic variation – NeoMundi AI Observatory
AI Observatory

11th AI Barometer – August 24 to 30, 2026

All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, AI Barometer #11 remains largely in the normal regime, with average stability virtually unchanged around 0.923. Global semantic variation reaches 1.93%, compared with 1.57% in Barometer #10. Factual alerts remain limited, and no generalized shift of the panel is observed. Reading:

What the Average Conceals - Analysis of AI Barometers #6 and #7 by James Moore - NeoMundi AI Observatory
AI Observatory

What a Stable Average Can Conceal in Production AI

James Moore’s analysis of AI Barometers #6 and #7 shows how a minimal global shift can mask concentrated changes across specific profiles, tasks and alert types. NeoMundi’s AI Barometers were designed to make the behaviour of AI systems observable over time. Once the measurement is produced, however, another question becomes essential: what can an apparently

AI Barometer #10 – Green level held, global stability – Fluent factual error undetected – NeoMundi AI Observatory
AI Observatory

10th AI Barometer – August 16 to 23, 2026

All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, AI Barometer #10 remains largely in the normal regime, with average stability virtually unchanged at 0.9229. Global semantic variation decreases to 1.57%, compared with 2.09% in Barometer #9, while factual alerts remain limited. The observed movement remains mainly localized on a

NeoMundi AI Cartography August 2026: 12 AI profiles plotted by OpenAI factuality (horizontal axis) and Mistral factuality (vertical axis), point size indicating mean observed stability
AI Observatory

AI Cartography of August 2026: Measuring the Behavior of 12 LLMs

Two systems with almost the same stability can differ by more than 30 points in factuality Protocol 1 – 12 systems × 790 TruthfulQA questions NeoMundi’s August monthly cartography once again compares twelve de-identified AI profiles across several complementary dimensions: their behavioral stability over 790 TruthfulQA questions, the factuality of their responses independently assessed by

AI Barometer #9 – Global stability, localized variation – Factual hallucination in a fluent response – Green level – NeoMundi AI Observatory
AI Observatory

9th AI Barometer – August 9 to 16, 2026

All AI Barometers Observation Level: 🟢 Green Across 4,800 executions distributed among 12 systems, Barometer #9 remains largely in the normal regime (97.12%), with average stability virtually unchanged at 0.9228. Semantic variation, however, rises to 2.06%, compared with approximately 1.06% in the 8th Barometer, while factual alerts remain limited. The observed movement is mainly localized

AI Barometer #8 – First time at green level, global calm confirmed – Collective stabilization of profiles – NeoMundi AI Observatory – Green level
AI Observatory

8th AI Barometer, week of August 3–9, 2026

All AI Barometers Observation Level: 🟢 Green This week, AI Barometer #8 confirms a return to global calm and moves to the green level. Two movements are particularly notable: This edition also presents a new example of a fragile factual claim embedded in a fluent and convincing response. This type of case, difficult to detect

AI Barometer #7 – Global calm confirmed, punctual signal at 9% – Week of July 27 to August 2, 2026 – NeoMundi AI Observatory – Yellow level
AI Observatory

7th AI Barometer, Week of July 27 to August 2, 2026

All AI Barometers Observation Level: 🟡 Yellow This week, AI Barometer #7 highlights two complementary dimensions. On the one hand, the overall calm observed in AI Barometer #6 is confirmed: average semantic variation decreases, and the majority of profiles remain close to their usual regime. On the other hand, this return to calm is not

Stability and Reproducibility of Runtime Signals – Fatima Ezzahrae Gouarab – NeoMundi Observatory
AI Observatory

NeoMundi publishes a first study on the stability and reproducibility of its runtime signals

Fatima Ezzahrae Gouarab, Data Scientist, statistician and scientific contributor to the NeoMundi Research Observatory, is conducting work on the stability and reproducibility of the signals produced by NeoMundi ControlTower. DOI : https://doi.org/10.5281/zenodo.21499715 This first experimental phase is based on 680 generations, carried out with three model providers, two generation temperatures and several corpora covering factual

Scroll to Top