How Galileo is checked
Galileo separates two questions: whether the available evidence supports action, and whether a report faithfully represents its computed outputs. The first is credibility assessment. The second is report integrity.
Credibility assessment
Can the evidence support the decision?
Galileo’s credibility assessment examines data sufficiency, channel separability, uncertainty, model diagnostics, material caveats, and the strength of the evidence behind a recommendation.
A strong historical fit does not automatically make a channel estimate identifiable or a budget recommendation actionable.
Report integrity
Does the report match its computed outputs?
Separate integrity checks verify relationships such as totals, interval ordering, response-curve labels, recommendation consistency, provenance, and bounded impacts. A failed check can prevent a report from being released.
These checks establish internal consistency. They do not establish that observational data has identified the true causal effect.
Synthetic ground truth
Recovery under specified synthetic conditions
On the always-on synthetic benchmark, effect-magnitude recovery remained within the benchmark’s gate, interval coverage was 100%, and interaction-detection precision was 1.00 against planted effects.
These results describe performance on that synthetic benchmark. They do not establish customer-data accuracy or universal estimator superiority.
The benchmark also shows an important limitation: with similar channel responses and roughly one year of weekly data, channel ranking can remain underpowered even when aggregate effect magnitude meets its gate.
Promotion handling remains an identified limitation. Under the current reference-centering treatment, adding promotion controls has not yet demonstrated improved channel-effect recovery on the relevant test. Galileo therefore makes no promotion-related accuracy claim.
Held-out prediction
A benchmark-specific error result
On a separate held-out synthetic benchmark, forecast error was approximately 11% MAPE, within the benchmark’s pinned 8–25% acceptance band.
This is a benchmark-specific predictive result. It is not a universal Galileo accuracy rate, a customer-validation result, or by itself a threshold proving that a budget recommendation is safe to act on.
Uncertainty boundary
What the published intervals include
The current published product artifacts report classical coefficient uncertainty. Some model-backed scenarios hold fitted saturation parameters fixed, so their intervals do not represent every source of model, causal, or decision uncertainty.
Agent provenance
Quantitative answers remain tied to computed evidence
Galileo’s agent layer is designed to allow quantitative claims only when they can be traced to model or scenario output. Requests for quantities absent from the analysis should produce a refusal rather than an estimate.
Updates and drift
New evidence can trigger a refit
When new weeks arrive, Galileo compares them with the expectations of the existing analysis. If the new data crosses the drift gate, it refuses an incremental update and calls for a full refit.
The public session replay demonstrates that refusal on synthetic data.