professionalinsight

Operationalizing Canada’s Federal Guideline OSFI E-23 — Model Risk Management to Deliver Fair Consumer Outcomes

By Frederick Au
O

ver the past several years, the CAS, through its research task forces, has extensively researched how various state and international regulators are approaching algorithmic fairness and model bias. As the global actuarial profession transitions from defining these frameworks to operationalizing them, Canada emerges as a live-environment test case. On May 1, 2027, the Canadian insurance industry enters a new era of governance. This date marks the deadline for full compliance with the Office of the Superintendent of Financial Institutions (OSFI) Guideline E-23 on Model Risk Management (MRM).1 While treating E-23 primarily as a rigorous federal compliance checklist is a defensible baseline for many institutions, integrating it with the broader market conduct goals creates the foundational infrastructure needed to navigate an environment increasingly scrutinized for algorithmic fairness, specifically the “fair consumer outcomes” mandated by regulators like the Financial Services Regulatory Authority of Ontario (FSRA).

We are entering a period where models, including those for insurance ratemaking and underwriting, should be mathematically sound, legally defensible, and socially fair. A model that is predictive but results in unexplained disparities is no longer just a market conduct issue; under the expanded scope of E-23, it may represent a model risk event or a compliance challenge.

The great convergence: A national imperative

For decades, the actuarial control cycle in Canada operated within a regulatory framework that treated financial risk and market conduct as separate domains. Historically, OSFI monitored whether the institution had sound risk management practices to maintain safety and soundness, while provincial regulators independently supervised whether the resulting rates and underwriting rules treated customers fairly. Under this bifurcated regime, the central question in model validation was often limited to: “Is this model predictive?” If a pricing model accurately predicted claims costs and secured the target equity returns, it was deemed a success, albeit with potential concerns on opacity or demographic impact.

Guideline E-23 alters this landscape by forcing these two worlds to interact. By expanding the definition of Model Risk to explicitly include adverse financial impact such as operational or reputational consequences,1 E-23 provides the governance chassis where these deliberate trade-offs are evaluated, documented, and justified by management.

A market-moving trend

While Ontario’s FSRA has been vocal with its proposed guidance on automobile insurance rating and underwriting,2 this convergence is driving the national agenda, led by Canada’s two largest regulators:

  • Ontario: FSRA’s guidance explicitly moves toward principles-based regulation, focusing on outcomes rather than technical rules.
  • Québec: The Autorité des marchés financiers (AMF) has released a guideline setting expectations for institutions to manage AI systems based on their impact on consumers.3

While the specific legal mechanisms differ among jurisdictions, E-23 provides the unified governance chassis to adapt to these evolving provincial expectations. Implementing a prudent E-23 MRM framework provides the evidentiary baseline required to demonstrate market conduct compliance to provincial regulators.

The legal landmine: The expiration of the “Zurich defense”

To understand the practical implications of this convergence, we should revisit the legal bedrock of Canadian actuarial practice: the Supreme Court of Canada’s 1992 decision in Zurich Insurance Co. v. Ontario.4 Under the Ontario Human Rights Code, insurers are permitted to use discriminatory rating variables only if the practice rests on “reasonable and bona fide” grounds. In Zurich, the Supreme Court established a rigorous two-part test to prove a pricing practice is “reasonable”: it must be based on a sound and accepted insurance practice (demonstrating a rational connection to the risk), and there must be no practical alternative. For 30 years, insurers have relied on this precedent to justify segmentation. However, the historical application of the “Zurich Defense” is facing re-evaluation driven by modern AI capabilities and stricter provincial oversight. Guideline E-23 accelerates this reckoning by mandating a risk-based approach to managing model risks that exposes whether a model truly possesses a rational connection or relies on discriminatory proxies.

Challenge 1: The rational connection (from correlation to causality)

In 1992, the Court accepted that a simple statistical correlation was sufficient to establish a rational connection. However, FSRA’s proposed new “Automobile Insurance Rating and Underwriting Guidance” fundamentally alters this standard such that statistical correlation is no longer a safe harbor if the variable acts as a proxy for a prohibited ground.2 In the age of AI, a model might find a statistical correlation between a permissible variable and a protected class. Under the old Zurich standard, the correlation might have been enough. Under FSRA’s fair consumer outcomes standard, this could be a direct or indirect proxy for unfair discrimination. Without the deep “explainability” required by E-23, an insurer cannot prove they are capturing a true risk driver rather than just a correlated bias.

For instance, in usage-based insurance, heavily penalizing late-night driving might correlate with the shift workers in lower-income brackets. Actuaries should consider using appropriate proxy variable tests to prove the risk lies in the fatigue and visibility of night driving, not the socioeconomic status of the driver.

Challenge 2: No practical alternative in the age of AI

The Supreme Court accepted the “no practical alternative” defense largely because the data required to price risk without using discriminatory proxies did not exist. Today, with E-23 mandating the identification of model limitations and mitigants, insurers face a higher evidentiary burden. They cannot simply assert that fairness is impossible; they must demonstrate it.
For instance, in usage-based insurance, heavily penalizing late-night driving might correlate with the shift workers in lower-income brackets. Actuaries should consider using appropriate proxy variable tests to prove the risk lies in the fatigue and visibility of night driving, not the socio-economic status of the driver.

A structured due diligence framework: The Human Rights Impact Assessment (HRIA)

This is where impact assessment tools, like the Human Rights Impact Assessment for AI (HRIA),5 developed by the Ontario Human Rights Commission (OHRC) and the Law Commission of Ontario (LCO), become critical. While the HRIA is a policy guideline tool rather than a binding legal shield, it provides a structured framework to document due diligence across both prongs of the Zurich test:

  1. Validating the Rational Connection: The HRIA advises insurers to evaluate statistical correlations, utilizing explainability tools to prove that variables are capturing genuine, causal risk drivers rather than acting as proxies for protected classes.
  2. Proving No Practical Alternative: If an adverse impact is identified, the HRIA recommends an alternatives analysis. By systematically testing less discriminatory models and generating privileged documentation that records the resulting degradation in predictive accuracy and financial viability, the HRIA establishes the evidentiary baseline required to debate “undue hardship” or lack of a commercially viable alternative before a regulator.

Integrating the HRIA into the E-23 validation process does not grant statutory immunity. However, it ensures that if an insurer retains a model with disparate impact, they do so with a documented defense that the model represents a sound insurance practice with no viable commercial or technical alternative.

Operationalizing E-23: Integrating model compliance risks into the model life cycle

Operating model compliance risk factors — such as significance of human impact, likelihood of harm, bias and fairness, and explainability — as a separate workstream from core MRM can create fragmented oversight. This siloed approach risks creating a blind spot where a model meets mathematical standards but presents potential legal or regulatory concerns. The statutory defense of a “reasonable and bona fide” practice might fail if an insurer cannot prove they rigorously assessed alternatives. Guideline E-23 serves as a mechanism for generating this proof by establishing a unified approach to compliance risk factors throughout the model life cycle.1

  • Risk Rating and Management Intensity: Insurers should establish a risk rating that moves beyond financial materiality to include key dimensions of compliance risk. For rating and underwriting applications, the significance of human impact, the likelihood of discriminatory harm, and the required level of explainability are critical factors in the inherent risk rating. These ratings drive the downstream model life cycle, determining model usage limits, monitoring intensity, and the escalation of residual risk management decisions.
  • Model Rationale and Documentation: Model owners should provide a clear rationale for deployment that explicitly addresses market conduct and fair consumer outcomes. This includes documenting considerations for the required level of transparency and explainability, as well as a proactive assessment of the potential for biased outcomes, negative social and ethical implications, or privacy risks.
  • Model Data and Development: The guideline expands data governance requirements from primarily accuracy concerns to broader facets: data should be relevant, representative, compliant, traceable, and timely. Insurers should enhance model explainability by analyzing the potential for unwanted data bias to translate into unfair model outputs and associated reputational risks. Clear, consistent, and repeatable practices for model development should be established to ensure that explainability standards are met, with rigor varying based on regulatory requirements and the potential impact on customers.
  • Model Review and Deployment: E-23 requires independent model review to confirm that the model outputs are appropriately explainable and comply with performance expectations before the model impacts a consumer. Crucially, deployment might necessitate conditional approval subject to outcome monitoring to detect whether “fairness drift” occurs post-launch, ensuring that the model remains fair not just in the test environment, but in the real world.

By operationalizing these E-23 principles, insurers can ensure that the necessary evidence for the “Zurich Defense,” i.e., the proof of diligence and the testing of alternatives, is sufficient and documented as part of the standard, enterprise-wide control cycle.

The E-23 Perimeter: A risk-based expansion beyond ratemaking and underwriting

Before tiering models, the primary operational hurdle for complying with E-23 is defining the model inventory. The guideline’s expanded definition captures everything from advanced machine learning algorithms to heuristic end-user computing tools. Applying maximum life cycle governance to every model would cause operational paralysis. Therefore, the E-23 blueprint should be applied through a risk-based approach that is proportional to the level of model risk identified by insurers.
High compliance risk models do not always calculate a premium; they can act as gatekeepers to the quoting process itself.

The regulatory dividend: Enterprise-wide confidence

While FSRA’s auto insurance guidance primarily targets rating and underwriting,2 OSFI E-23 mandates an expectation of enterprise-wide coverage, subject to materiality. By applying E-23 rigor to models outside the strict scope of pricing and underwriting, insurers provide provincial regulators with a higher level of confidence that consumer fairness is being managed holistically across the value chain.

The risk-based expansion can be illustrated through three tiers of operational reality:

  • High compliance risk models do not always calculate a premium; they can act as gatekeepers to the quoting process itself. Consider an algorithmic point-of-sale fraud model that evaluates a digital footprint. If an applicant is scored as “high risk,” the system intentionally injects quoting friction, such as blocking the direct-to-consumer online rate and forcing a manual broker call. If this model relies on proxy variables that systematically flag specific minority cohorts, it could constitute a discriminatory barrier to entry for a mandatory financial product. Because these models dictate fundamental, equitable access to coverage, those resulting in systematic, disparate barriers require a full “reasonable and bona fide” assessment. Insurers should use human impact assessment tools like the HRIA to prove the fraud variables capture genuine, causal risk rather than acting as protected-class proxies and explicitly demonstrate a lack of less discriminatory screening alternatives.
  • Medium compliance risk models prioritize convenience, creating an indirect fairness impact that requires lighter control. For example, a claims triage model that decides who gets instant approval versus standard handling creates a conduct risk if one group is systematically slowed down, but it does not accuse the customer of fraud. While these models may not demand an exhaustive assessment, they need sufficient pre-deployment proxy testing on historical data combined with automated post-deployment circuit breakers to ensure service level disparities remain within acceptable bounds.
  • Low compliance risk models have remote or nonexistent human impact. Applying fairness testing here would be a misuse of resources. For example, actuarial reserving models operate on aggregate data pools to ensure solvency. While crucial for financial stability, they do not make individual decisions about consumers. For these models, impact assessment tools like the HRIA are non-applicable. The focus remains on the traditional pillars of performance and stability. By explicitly categorizing these as low compliance risk that are subject only to light inventory requirements, the insurer demonstrates the “proportionality” required by OSFI, preserving resources for the highest impact models.

The path forward: Operationalizing E-23 to deliver fair consumer outcomes

While achieving compliance with the OSFI E-23 operational deadline is the baseline objective, its implementation and integration with broader market conduct goals offer a distinct advantage. In jurisdictions like Ontario, insurers who leverage E-23 to build a fair modeling ecosystem can position themselves favorably for supervisory proportionality, potentially achieving greater regulatory efficiency and speed to market for their rate filings.
To navigate this successfully, the industry should focus on:

  1. Integrated life cycle management: The end-to-end model life cycle should explicitly integrate model compliance parameters for fair consumer outcomes.
  2. Risk-based governance: Governance rigor should be proportional to the model compliance risk parameters such as bias, fairness, explainability, and human impact.
  3. Evidentiary escalation versus risk acceptance: Market conduct violations cannot be formally accepted like financial and insurance risks. Models exhibiting unmitigated disparate impact should be escalated to senior management and legal counsel strictly to validate the “Zurich defense” prior to deployment.

The convergence of E-23 and FSRA requirements on fair consumer outcomes represents the current trajectory. Actuaries should review their model inventories not just for financial materiality, but also for compliance materiality. As the industry transitions into a regulatory environment that demands higher transparency, proactively operationalizing risk-based fairness provides the essential infrastructure to navigate these evolving standards effectively.

Fredrick Au, FCAS, is a member of the CAS Canada Race & Insurance Pricing Research Task Force and an actuary at TD Bank.

References

  1. Office of the Superintendent of Financial Institutions (OSFI). Guideline E-23: Model Risk Management (2027)
  2. Financial Services Regulatory Authority of Ontario (FSRA). Guidance: Automobile Insurance Rating and Underwriting Supervision (No. AU0142INT)
  3. Autorité des marchés financiers (AMF). Guideline for the Use of Artificial Intelligence. June 2025.
  4. Zurich Insurance Co. v. Ontario (Human Rights Commission), [1992] 2 S.C.R. 321.
  5. Law Commission of Ontario & Ontario Human Rights Commission. Human Rights Impact Assessment (HRIA) for AI. November 2024.