If an AI system changes its behavior without changing its name, how would we notice?
A single response is not enough. The same system can produce different answers to an identical question. Its behavior can also evolve over time, particularly when its model, settings, or infrastructure change.
AI Weather™ makes some of these variations visible. Each day, it observes AI systems under a repeated protocol and publishes an observed behavioral condition alongside the underlying measurement data.
That condition is neither an overall grade nor a ranking of models. It is a signal at a specific point in time, produced under defined conditions.

The principle: keep a point of comparison
To tell whether something has changed, we need a reference.
AI Weather uses fixed sentinel questions submitted regularly to the systems it observes. Repeated runs reveal how much responses vary within a day. Successive daily observations make it possible to compare that behavior with the same system’s recent history.
The condition shown by AI Weather comes from this longitudinal channel: comparable prompts observed over time. A new question may also be posed in “Today’s Challenge” to explore a different behavior. That is a separate experiment and does not determine the AI Weather condition.
This separation matters. A new question may produce an interesting, surprising, or problematic answer. On its own, it cannot establish that a system’s usual behavior has changed.
How is a daily condition produced?
The protocol has four steps.
1. Observe. Systems in the panel receive sentinel questions under a documented protocol. Their responses are recorded with the context and time of execution.
2. Compare. The day’s observations are compared with the same system’s history. The measurement examines factors including stability, the size of variations, and how they develop over time.
3. Assess the strength of the signal. A one-off difference, a persistent change, and a day with missing observations should not be interpreted in the same way. Persistence and protocol coverage therefore contribute to the published condition.
4. Publish. Results are made available as dated data, so that the displayed condition can be traced back to the observations and the version of the method that produced it.
A color makes it easier to see where attention may be needed. It does not replace examination of the data or human judgment.
What exactly is being measured?
AI Weather observes a system’s behavior under a specific protocol. It aims to make regularity, variation, and potential changes in behavioral patterns visible over time.
Stability does not mean that answers are true: a system can consistently repeat an error. Conversely, a different answer is not necessarily a bad one. The AI Weather condition therefore establishes neither the truth of each response nor the overall quality of a model.
This distinction guides NeoMundi’s work on measurement: defining what is observed, specifying how each metric is produced, stating when it can be used, and documenting its limits. A number becomes useful only when we can understand what it measures and under what conditions we can trust it.
An instrument being built in public
Since June 1, 2026, NeoMundi has tested several ways to observe AI behavior, including barometers, behavioral maps, and AI Weather. Together, these initiatives have produced more than 310,000 observations across the NeoMundi program. This figure covers the program as a whole; it does not refer solely to AI Weather’s daily protocol.
Accumulating data is not enough to establish the quality of an instrument. Protocols must also be audited, metrics defined, results tested against their limitations, and methods improved when experience reveals a problem.
That is why AI Weather is presented as an experimental instrument open to examination. Its results can be studied, challenged, and tested. AI Weather’s public code and resources are available on GitHub.
Why does this experiment matter beyond AI Weather?
AI Weather is a public demonstration of a broader need: between AI systems and the organizations that use them, there is often a missing independent layer for measuring actual behavior.
NeoMundi is developing that layer. Its role is to produce a traceable signal during execution and over time. Depending on the context, that signal can support analysis, an audit, an application, or governance rules set by the organization.
The value of such a measurement depends on how it is used: understanding a variation, documenting an incident, comparing two periods, or deciding when human oversight is needed. NeoMundi provides the signal; users interpret it and decide what to do with it.
What AI Weather cannot tell us
Thirty daily runs, or any other limited set of observations, cannot summarize an AI system’s overall behavior. Observations based on sentinel questions cannot guarantee how the system will behave with every question, in every context, or the next day.
AI Weather does not identify the “best” model. It does not certify any provider, predict future responses, or replace an audit tailored to a specific use case or a compliance assessment.
It asks a narrower, testable question: under this protocol, is the behavior observed today still comparable to that of previous days?
That is a starting point for journalists, researchers, data scientists, auditors, and organizations that want to examine the data, test the method, or propose a better way to measure.
Measure before interpreting. Observe before governing.
