Plain-language definitions of the terms we use in model validation, AI assurance and algorithmic risk review.
Independent technical review of a model's performance, calibration, robustness, drift, fairness and operational risks before it is trusted in a real workflow.
A structured process for checking whether a model performs as expected on relevant data and within its intended use.
Independent review of automated decision systems, including scoring models, rules, workflows and AI systems that affect high-impact decisions.
A change in data, behaviour or environment that makes a model less reliable over time.
A measure of whether predicted probabilities match real outcomes. If a model says 80%, it should be right about 80% of the time.
Assessment of whether a model behaves differently across relevant groups or creates unequal error patterns.
A model's ability to remain reliable under noise, edge cases, perturbations or changing conditions.
When a model or strategy performs well on past data but fails on new data because it learned noise instead of signal.
Testing and monitoring of LLM agents, copilots and generative AI systems for hallucinations, jailbreaks, prompt injection, unsafe outputs and production failure modes.
Evaluation of whether an LLM produces confident but false, unsupported or misleading outputs.
An attack or failure mode where external or user-provided text manipulates an LLM into ignoring instructions or leaking unsafe behaviour.
Testing whether an LLM can be pushed into bypassing safety rules, business rules or expected refusal behaviour.
The governance process used to identify, measure, monitor and control risks created by models.
An AI system that may fall under high-risk categories defined by the EU AI Act. Model Assurance Lab does not certify compliance; it provides technical evidence that can support governance and legal review.
A 0–100 technical score summarizing evidence strength, model reliability and operational risk signals within the reviewed scope.
When a trading strategy looks profitable historically because it was tuned too closely to past data.
A validation technique that repeatedly tests a strategy on future unseen periods after training or tuning on earlier data.
A method for testing robustness by repeatedly reshuffling or simulating outcomes to estimate uncertainty, drawdown risk or dependence on luck.
Need a term explained for your model or strategy? Request a review.