The Classical Model Validation Process
Foundations, Limitations, and Perspectives
The validation of risk models is a core component of risk management in banks and is shaped by strict regulatory requirements. At the same time, there is no uniform standard, as approaches vary depending on internal processes, technical capabilities, and model-specific characteristics.
We begin by providing an overview of the classical validation process. Traditional validation methods are increasingly reaching their limits in a data-driven and dynamic environment, whether due to the growing complexity of modern models, rising transparency requirements, or the need for more efficient automation.
Building on this, we provide an outlook on further discussions around modern validation approaches, including automation solutions and the use of AI challenger models, which open up new potential in model review. At this stage, however, the focus remains on the classical validation process, which serves as the starting point for the further development of methods and technologies.
Steps in the Validation Process
The first step in the classical validation process is to ensure data integrity, which forms a key basis for valid model predictions. The following aspects are the main focus:
Data Quality: The data are reviewed for accuracy, consistency, and reliability to ensure that they are suitable for modeling.
Representativeness: An assessment is made as to whether the data adequately reflect the target population or target process.
Completeness: Missing values are identified and, where possible, supplemented using appropriate techniques such as imputation.
Outliers: Potential outliers that could impair model quality are identified and treated accordingly.
Bias and Fairness: The data are examined for potential biases in order to avoid discriminatory patterns. Fairness analyses complement this step, particularly for models involving potentially critical decisions.
For the use of the data, they are further prepared in a structured process:
Filtering: Relevant data records are identified and extracted.
Transformation: The data are converted into a format that meets the model requirements, for example through normalization or encoding.
Cleansing: Erroneous or redundant data are removed to ensure model stability.
In the second step of the validation process, the input data are prepared by the data processor for model assessment. The objective is to divide the data meaningfully into training and test data in order to enable an objective assessment of model quality. In classical validation, specific approaches such as backtesting or year-on-year comparisons are often used.
This intermediate step provides the basis for a meaningful and transparent model assessment and ensures that the validation results are robust and practically relevant.
Partitioning into Training and Test Data:
The input data are divided into two or more data sets:
- Training Data: Used to train the model and calibrate its parameters.
- Test Data: Left untouched to ensure an independent review of model quality.
Care is taken to ensure that the data distribution, for example class ratios, remains representative in both data sets.
Backtesting:
The model is tested on historical data in order to assess how well it would have predicted past events. This provides insight into the model’s stability and accuracy under realistic conditions.
Year-on-Year Comparisons:
Especially in time series models or in scenarios involving changing data, it is examined whether the model maintains its predictive power over multiple periods.
Once the data have been prepared in the data processor and divided into training and test sets, model training and the subsequent test phase are carried out in order to assess the model’s performance. This step is critical to ensure that the model is both robust and suitable for practical use.
Model Training (Train):
Using the training data, the model is trained to identify patterns in the data and generate predictions. The objective is to optimize the model parameters so that predictions remain both accurate and stable.
Control via Hyperparameters:
Hyperparameters, such as thresholds (theta), are set to adapt model performance to specific requirements. Thresholds can be used to control the sensitivity of the model, for example to define which predictions are classified as risk. Hyperparameter selection is often performed using techniques such as grid search or random search in order to find the optimal combination.
Test Phase with Forecasts (Predict):
After training, the model is evaluated using the Testdaten, which were not used during training. The objective is to generate realistic forecasts (predict) that assess the model’s practical suitability and robustness.
Assessment and Evaluation of Results
Result Metrics:
The model results are assessed using specific metrics. These may include score values, probabilities of default, or expected losses.
P&L Distribution and Derived Risk Metrics:
In addition to classical model metrics, the profit and loss distribution (Profit and Loss) is often analyzed. Risk metrics such as Value-at-Risk (VaR) are generally derived from this distribution. It shows how the model affects decisions in terms of financial gains and losses as well as risk management and provides insight into the model’s economic robustness. A detailed analysis of the P&L distribution and the associated quantile information makes it possible to identify weaknesses that could impair the model’s stability or economic viability.
Following training and the initial test phase, the validation process proceeds with a detailed assessment of model quality to ensure that the model is not only accurate, but also robust and stable. The results of this assessment form the basis for deciding whether a model can be transferred into production or whether it must be further optimized. In the diagram, this step is referred to as “model quality” and is decisive in evaluating the model’s reliability and performance.
Assessment of Model Quality:
Model quality is a collective term describing the overall quality of a model. It assesses how well a model fulfills its intended purpose and how much predictive power it has. In doing so, model quality considers several dimensions:
Robustness: Robustness describes a model’s ability to make correct predictions even under uncertain or changing conditions. It concerns how well the model responds to deviating or erroneous input data or unforeseen scenarios.
Stability: Stability describes a model’s ability to deliver consistent results over time and across repeated applications. It concerns the reliability of predictions under unchanged conditions, that is, how sensitive the model is to small changes in input data or model structure.
Accuracy: The accuracy of a model indicates how many predictions, for example default or non-default, are correct, measured as the proportion of correctly classified events relative to the total number of predictions.
Calibration: Calibration assesses how well the predicted probabilities align with actual outcomes. It is particularly important in the practical application of probabilistic models.
Confidence Interval: Confidence intervals quantify the uncertainty in a model’s predictions or parameters. They indicate the range within which the actual values lie with a given probability.
Significance Testing:
Significance is a statistical concept and is particularly relevant in model validation when assessing the reliability of results and estimates. Significance measures whether an observed effect or relationship in the data is unlikely to have arisen by chance. It is especially important in the assessment of models, parameters, and hypotheses. Significance plays a specific role in validation and is clearly distinct from concepts such as accuracy, calibration, or robustness. It complements these concepts by evaluating the statistical reliability of model statements and characteristics. Typical tests, such as the Kolmogorov-Smirnov test, include p-values or hypothesis tests for validating model parameters.
Following the assessment of model quality, a detailed analysis block is carried out in which the resilience and consistency of the model are comprehensively reviewed. Unlike model quality tests, which focus on performance and the fundamental quality of a model, analyses generally aim to generate deeper insights into the model’s functioning, structural validity, and practical implications. They complement the assessment of model quality at a more granular level. The objective is not only to further evaluate the quality of model predictions, but also to make the underlying assumptions and the contribution of individual variables transparent.
Consistency Checks:
The objective of consistency checks is to determine whether the model produces consistent results under identical or similar conditions. Consistency checks are closely linked to model stability, which has already been examined during the model quality assessment. However, consistency checks are more detailed and are designed to uncover specific inconsistencies in the results or in the underlying processes, such as data preparation.
Materiality Assessments:
Materiality assessments are used to determine how strongly individual model components or variables influence the results. Unlike general model quality tests, which measure overall performance, materiality assessments focus on the relevance of individual components. They help identify which parameters or variables are decisive for model predictions and make it possible to identify or eliminate less significant factors.
Attribution and Contribution Analyses:
These analyses make it possible to identify and quantify the influence of individual variables (attribution) or groups of variables (contribution) on model predictions. They go beyond model quality by making the model logic transparent. In this way, they provide insight into how variables interact and whether their effects are consistent with the model assumptions. While attribution analyses assess specific variables, contribution analyses consider larger groups or categories, which makes them particularly relevant for complex models.
Stress and Sensitivity Tests:
Assessing a model’s robustness by simulating extreme scenarios (stress) or by selectively varying individual input factors (sensitivity) focuses on model resilience and complements the robustness tests within model quality assessment. However, these tests place a stronger emphasis on practical scenarios and regulatory requirements by examining how the model responds under rare or extreme conditions
The final step in the validation process is the preparation of the validation report, which summarizes and documents the results of the entire review. This step serves not only internal transparency, but also forms a key basis for communication with regulatory authorities and other stakeholders.
Documentation:
All tests, analyses, and methods carried out are documented in detail, including the data, models, and parameters used. Documentation is essential to ensure the traceability of the validation process and to meet regulatory requirements.
Evaluation of Results:
All findings from the previous process steps are systematically consolidated. The evaluation includes quantitative results, such as model quality, calibration, and robustness, as well as qualitative analyses, such as attribution and sensitivity analyses. Particular attention is paid to the identification of weaknesses or risks uncovered during validation.
Recommendations and Actions:
Based on the results, the report provides clear recommendations for improvements or adjustments to the model. These may include, for example, adjusting hyperparameters, extending the data set, or performing additional tests. In the case of material weaknesses, specific action plans are developed to address them.
Margin of Conservatism (MoC):
The Margin of Conservatism represents a safety add-on or conservative adjustment applied in risk models to compensate for uncertainty and weaknesses in the modeling. The MoC is applied to address inherent model risk arising from uncertainty in assumptions, data, or methodologies. It ensures that risk assessments are sufficiently conservative to avoid underestimating the actual risk. The goal is to ensure that banks hold enough capital to cover losses even under real and stressed market conditions.
Classical validation provides a solid foundation for model review. It establishes the basis for robust risk management and ensures that regulatory requirements are met. At the same time, advances in AI and machine learning are opening up new opportunities to complement classical procedures and overcome their limitation.
The classical validation approach has proven effective over the years and meets the fundamental requirements of model assessment in banks. Nevertheless, this approach is increasingly reaching its limits, particularly in a more complex, data-intensive, and dynamic world. The typical challenges can be summarized as follows:
- Limited Flexibility for Complex Models:
Classical approaches are often based on linear or statistical models that are only able to capture the high-dimensional and non-linear relationships in modern data to a limited extent. Models such as Monte Carlo simulations are occasionally used as challenger approaches, but they provide only limited coverage of variability and uncertainty. - High Manual Effort:
Data preparation (data integrity), testing, and documentation are often time- and resource-intensive. Especially with large and heterogeneous data sets, this requires substantial capacity and extends validation cycles. - Limitations in the Analysis of Results:
Classical methods provide clear output metrics, such as discriminatory power or calibration, but they reach their limits when it comes to uncovering complex interactions between variables or potential weaknesses in model behavior. - Static Review Processes:
Classical validations are often performed in fixed cycles, such as annual validation, making it difficult to respond promptly to dynamic changes in the data or underlying conditions. - Lack of Automation:
The absence of automation in many steps, such as attribution or sensitivity analysis, makes efficient model assessment more difficult and reduces the possibilities for continuous monitoring. - Limited Use of Classical Challenger Models:
Although challenger models such as Monte Carlo simulations or simple scenario tests are occasionally used, they often lack the ability to identify deeper patterns and non-linear relationships in the data. Their function is generally limited to confirming or rejecting individual model assumptions.
While classical validation approaches provide a solid foundation, they often cannot keep pace with the requirements of modern risk models. This is where AI-supported challenger models open up entirely new opportunities to extend and optimize existing validation processes:
Enhanced Pattern Recognition: AI can detect deeper relationships in the data that remain hidden to classical models, thereby providing new perspectives on model assumptions and outcomes.
Automated Monitoring: AI enables continuous real-time monitoring and assessment of models, which is particularly important in rapidly changing market environments.
Greater Efficiency: Through the use of AI, time-consuming steps such as sensitivity analyses or documentation can be partially or fully automated.
More Comprehensive Challenger Approaches: AI-based challenger models expand the possibilities of not only reviewing existing approaches, but actively improving them.
In a subsequent article, we will describe in detail how AI challenger models can complement classical validation processes, which prerequisites must be met for their successful use, and which regulatory aspects need to be taken into account.
As a specialized consulting firm, we support banks in the regulatory-compliant implementation of ML models, from strategy development through to sustainable integration into existing processes.
Would you like to learn more about AI-supported model validation? Contact us for a non-binding consultation.


