The most common framing of AI ethics in Canadian public discourse goes roughly like this: AI systems sometimes produce biased outputs, the bias reflects bias in training data, the solution is to clean the data and audit the models, and once the technical problems are fixed, AI will be fair. The framing has the structure of an engineering problem with an engineering solution. This chapter argues that the framing is wrong, or more precisely that it's incomplete in ways that matter in practice. The bias-and-fairness conversation in AI is real, the technical problems documented in the academic literature are real, and the proposed fixes have genuinely improved things in some contexts. But the conversation, taken alone, treats the symptoms while leaving the structure that produces them intact. The deeper question, surfaced most clearly by the Lewis, Whaanga, and Yolgörmez Abundant Intelligences paper introduced in Chapter 1, is not whether AI can be made less biased within its current foundations. It is whether the foundations themselves are doing what they are claimed to be doing, and whose interests they were built to serve. This is the chapter's central methodological move. AI ethics is not primarily an ethics problem; it is an epistemology problem. Once that move lands, the technical bias-and-fairness literature reads as one layer of a larger conversation and the corporate ethics-washing patterns become easier to spot. And the Indigenous-led frameworks introduced throughout the guide come into focus as the alternative they actually are, rather than the diversity-checkbox they're sometimes treated as. You will leave with: the working distinction between the dominant Western corporate AI ethics paradigm and the alternatives that engage epistemology directly; specific documented Canadian cases of algorithmic bias and their lessons; the limits of technical fixes; the algorithmic fairness literature's genuine debates honestly presented; the connection between bias and the power-structure analysis the guide has been building throughout; the Abundant Intelligences constructive program engaged in practice; and the working test for evaluating any AI ethics claim. ---
What "bias" actually means in AI systems
Start with the term itself, because public discussion uses it in at least three distinct senses.
Statistical bias. The technical sense from the underlying mathematics: a systematic deviation between an estimator's expected value and the true value of the parameter being estimated. In machine learning specifically, statistical bias refers to errors that reflect the model's assumptions rather than the data's randomness. A system trained on a dataset that contains predominantly male faces will have lower accuracy on female faces; the lower accuracy is statistical bias in the technical sense.
Social bias. The everyday sense: systematic favouritism for or against specific groups of people based on demographic characteristics. A hiring AI that consistently rejects applications from racialized candidates exhibits social bias regardless of whether the underlying mathematics shows statistical bias against a specific parameter.
Structural bias. The sense the guide has been developing through Chapters 1, 3, 9, and 10: the way AI systems reproduce and amplify the power structures and resource distributions of the societies that produce them. The training data over-represents some voices and under-represents others. The labelling labour is performed by low-wage workers in the Global South. The deployment decisions are made by the corporate actors with most direct economic interest. The benefits accrue disproportionately to investors and to specific user populations. Each of these is bias in the structural sense even when no specific algorithm in the system shows technically measurable bias against any specific protected group.
These three senses are related but not identical. A system can have low statistical bias and high social bias (if the underlying training data accurately reflects discriminatory historical patterns, technically-unbiased prediction reproduces the discrimination). A system can have low social bias on measured protected categories and high structural bias (if it doesn't disadvantage any specific demographic group but still concentrates economic value with a small group of corporate actors). The conversations about AI bias often slide between these senses in ways that produce confused arguments.
The guide asks readers to track which sense is being used in any specific bias claim. Most corporate AI ethics statements address statistical bias and (to a lesser extent) social bias. They rarely address structural bias because structural bias would require structural changes that the corporate framework isn't designed to produce.
---
Documented Canadian cases
The bias conversation often happens at high abstraction. Grounding it in specific Canadian cases provides the empirical baseline for what the conversation is actually about.
Clearview AI and Canadian law enforcement. Discussed in Chapter 10 as a privacy case, the Clearview AI matter is also a bias-and-fairness case. Independent academic research has documented substantially higher false-match rates for facial recognition AI on darker-skinned individuals, on women, and on younger people. The Gender Shades study by Joy Buolamwini and Timnit Gebru (MIT, 2018) found commercial facial recognition systems with false-positive rates more than ten times higher for darker-skinned women than for lighter-skinned men. Subsequent NIST evaluations (National Institute of Standards and Technology, 2019 and updates) have confirmed and characterized these disparities across most commercial facial recognition systems. When Canadian law enforcement agencies used Clearview AI without proper authorization, the disparate accuracy rates meant that racialized Canadians, particularly Black and Indigenous Canadians, were disproportionately exposed to wrongful identification. This is a documented bias pattern in a documented Canadian deployment, not an abstract concern.
GOBLIN FACTS — "more than ten times" has actual numbers. In the 2018 Gender Shades study, commercial gender-classification systems misread darker-skinned women up to 34.7% of the time and lighter-skinned men 0.8% of the time (Buolamwini & Gebru). Same product, same demo. One number for the brochure, another for the people it works worst on.
🧌 GOBLIN CHECK — A system that's 99% accurate on average and 90% accurate on your face is not, for you, 99% accurate. Averages are where disparities go to do public relations. Whenever someone quotes one accuracy number for a system used on millions of different faces, ask for the breakdown. Buolamwini and Gebru asked for the breakdown, and the breakdown was the whole story.
Toronto Police's risk-prediction tools. The Toronto Police Service has explored multiple algorithmic risk-prediction tools over the past decade — software that aims to predict where crime is likely to occur, which individuals are likely to reoffend, and which calls require what kind of response. These tools have been the subject of substantial criticism from civil liberties organizations, the Canadian Civil Liberties Association, and racialized community advocates on the grounds that they reproduce existing patterns of policing, which themselves reflect documented histories of racialized over-policing. Toronto's specific deployments and their outcomes are partially documented; comprehensive independent audits are limited. The structural bias concern is that algorithmic risk prediction trained on historical policing data will direct future policing toward the same communities that were historically over-policed, producing a feedback loop that the algorithm experiences as accurate prediction.
Federal benefits and eligibility algorithms. Service Canada and provincial benefit administration systems use increasingly automated processing for employment insurance eligibility, social assistance determinations, immigration status decisions, and other consequential administrative actions. The Treasury Board AI Register (Chapter 10) documents many of these systems. Several have been the subject of legal challenges or administrative complaints, instances where the algorithmic processing produced outcomes that the underlying statute's authors had not contemplated. The Supreme Court's 2018 decision in Ewert v. Canada (treated in full in Chapter 10) remains the closest thing Canadian law has to a leading case on bias in algorithmic-style decision tools, and it predates the current AI wave. The broader pattern: federal and provincial governments deploy AI in decisions affecting Canadians' basic eligibility for benefits, status, and services, and the bias-and-fairness implications of those deployments are visible primarily through individual challenges rather than through systematic public audit.
Healthcare AI deployments. Several Canadian health authorities deploy AI in diagnostic imaging, triage decisions, predictive analytics, and resource allocation. The benefits in some applications (imaging accuracy for specific cancers, sepsis prediction, mental health screening) are real and documented. The bias concerns are also real: training data that under-represents specific populations (particularly Indigenous communities, racialized communities, and certain age and sex demographics) produces systems with lower accuracy on those populations. **The classic 2019 Science paper by Obermeyer and colleagues documented a US healthcare risk-prediction algorithm that systematically under-allocated care to Black patients because it used historical healthcare spending as a proxy for healthcare need, and historical healthcare spending was lower for Black patients due to documented disparities in access.** The algorithm wasn't biased against Black patients in any direct demographic sense; it was biased against them through the structural choice of which proxy to use. The same pattern almost certainly operates in Canadian healthcare AI but is less publicly documented because Canadian healthcare data is less accessible to independent researchers.
Hiring and employment screening. Documented in Chapter 10's workplace surveillance discussion, AI-driven hiring tools have been shown across multiple studies to exhibit bias against women, racialized candidates, older workers, candidates with disabilities, and those with non-Western names. Amazon famously abandoned its internal hiring AI in 2018 after discovering it had learned to disadvantage applications mentioning women's colleges and certain women's professional associations. **Ontario's Working for Workers Act (effective January 1, 2026) requires employers with 25+ employees to disclose AI use in hiring screening but does not require bias audits or limit the use itself.** The disclosure-without-restriction model is the working baseline in Canada.
These cases are not exceptional examples chosen to indict AI. They are typical examples of how AI deployments work when bias-and-fairness considerations are addressed through corporate ethics frameworks rather than through structural protections. The pattern is documented across hundreds of additional cases globally. The bias question is not whether documented bias cases exist; the bias question is what the structural response to documented bias should be.
Canadian public response to these documented patterns is stronger than the policy response would suggest. The Office of the Privacy Commissioner's 2024-25 privacy opinion research found that 88% of Canadians reported some concern about their personal information being used to train AI systems, including 42% who reported being extremely concerned. Trust in the entities deploying AI is also low: only 28% of Canadians have a fair amount or great deal of trust in "Big Tech" to protect personal information, and only 12% have equivalent trust in social media companies. These numbers bear on the public-confidence dimension of the bias-and-fairness conversation: when documented bias cases compound with low institutional trust, the legitimacy gap that corporate ethics frameworks are meant to address grows rather than narrows. Whether the gap can be closed through more elaborate ethics statements, or whether closing it requires the structural changes the rest of this chapter engages, is the methodological question the guide asks readers to hold.
---
The algorithmic fairness literature — what it offers and what it doesn't
The academic field of algorithmic fairness has produced substantial technical and conceptual work over the past decade. Honest engagement with the field requires recognizing both what it has accomplished and what its limits are.
What the field has accomplished. Multiple formal definitions of fairness have been developed and put to use: demographic parity (equal outcome rates across groups), equalized odds (equal error rates across groups), calibration (predicted probabilities match actual rates across groups), individual fairness (similar individuals receive similar predictions). Tools and techniques exist for measuring bias in deployed systems, for adjusting training procedures to reduce specific bias metrics, and for documenting model performance across demographic categories. Major AI labs and many enterprise deployments now include some form of bias-and-fairness testing as standard practice. The work has produced real improvements in specific contexts — facial recognition systems' disparate accuracy rates have narrowed substantially since the 2018 Gender Shades findings, healthcare AI tools have started using more representative training data, and hiring tools have been audited and adjusted.
The intellectual contribution. The 2016 paper by Kleinberg, Mullainathan, and Raghavan, and subsequent work by Chouldechova and others, established a fundamental result: multiple intuitive definitions of fairness are mathematically incompatible with each other in many real-world scenarios. A system that satisfies demographic parity may fail to satisfy equalized odds, and vice versa. A system calibrated equally across groups may exhibit different error rates across groups. The impossibility result means that "make AI fair" is not a well-formed engineering goal until the specific definition of fairness is chosen, and the choice involves value trade-offs that are political and ethical rather than purely technical. This is important intellectual work and the field deserves credit for it.
What the field's limits are. Three structural limits worth naming:
First, the focus on outputs. Algorithmic fairness research focuses primarily on the outputs of AI systems — whether the predictions are equally accurate across groups, whether the decisions are equally favourable. It addresses less directly the upstream questions of who built the system, who chose the objective, who funded the deployment, and who benefits from the deployment. A perfectly demographically-parity-compliant hiring AI deployed in a corporate environment where racialized executives are underrepresented in the hiring decision will still tend to reinforce that underrepresentation, because the metric of "good hire" was defined by the existing executive population. The fairness fix at the output layer doesn't address the bias at the metric-definition layer.
Second, the focus on protected categories. The legal infrastructure of anti-discrimination law (in Canada, the Charter of Rights and Freedoms section 15, the Canadian Human Rights Act, provincial human rights codes) provides a list of protected categories. Algorithmic fairness research has primarily organized around these categories. The structural biases that don't map onto protected categories (economic bias, geographic bias, language bias, generational bias, network-effect bias) receive less attention. A system that disadvantages rural Canadians, low-income Canadians, francophones, or workers in declining industries may not be addressable through the algorithmic fairness framework even though the disadvantage is documented and significant.
EXAMPLE — the postal code that means race. Tell a model to ignore race and it will reach for a stand-in: the postal code, the surname, the high school. Strip those and it finds new ones. This is why "we removed the sensitive field" so rarely removes the bias — the world is correlated, and the model is exceptionally good at noticing.
Third, the focus on technical solutions. Even when the fairness literature names structural problems, the proposed solutions tend toward technical fixes: cleaner data, sharper metrics, better-audited algorithms. The solutions that would require structural changes (different ownership of AI systems, different funding and governance models, different decisions about which problems get solved at all) receive less attention, because they sit outside the field's methodological frame. This is not a failure of individual researchers. It's a structural feature of how academic fields draw the boundary around what they study.
Taking stock. The algorithmic fairness literature is doing real work, producing real improvements, and developing genuinely important intellectual contributions. It is also operating within a frame that addresses some of the questions raised by AI bias-and-fairness concerns and not others. Reading the field as the full conversation about AI ethics, rather than as one layer of a larger conversation, produces a partial picture that the field itself sometimes encourages. The deeper questions about whose interests AI serves and how those interests get encoded into systems are not primarily questions the algorithmic fairness literature is designed to answer.
---
Disability — the dimension the fairness debate keeps missing
There is a group the algorithmic-fairness literature has the hardest time seeing, and it is the group AI most often fails: disabled people. Both halves of that sentence are true at once, and holding them together is the point.
Start with the genuine benefit, because there is real benefit. For many disabled people, AI is the most useful assistive technology in a generation: live captioning for Deaf and hard-of-hearing users, speech-to-text and text-to-speech, real-time image descriptions that let a blind person "read" a photograph, and communication tools for people who do not speak. These are not hypothetical. They are in daily use, and they expand access in ways earlier technology could not. A guide that credits AI's capabilities where they are real has to credit these.
Now the failures, which are structural rather than incidental. AI systems are trained to perform well on the average case, and disability is, almost by definition, the set of bodies and voices that sit at the edges of that average. Speech recognition that works for typical speech fails for atypical speech (cerebral palsy, ALS, a stammer). Facial analysis built on typical faces misreads facial differences, and "engagement detection" that scores eye contact penalizes autistic and blind users for whom eye contact is not the signal it assumes. Exam-proctoring AI deployed during the pandemic repeatedly flagged disabled students as "suspicious" for looking away, for moving differently, for using assistive tools. Video-interview hiring systems score candidates on expressions and vocal patterns that disabled applicants may not produce. And automated benefits systems, the software that helps decide eligibility for disability support and home care, have in documented US cases cut hours from disabled people through opaque formulas that no caseworker could explain. Canada automates more of this every year. The documented Canadian cases are thinner, which (as with healthcare in Section Two) reflects less access for researchers, not less risk.
🧌 GOBLIN CHECK — When a vendor says a system is "99% accurate," the goblin's first question is for whom? Accuracy is an average, and disabled people are exactly the bodies an average is built by setting aside. "Works for most people" is no reassurance to the person it doesn't work for. It is a description of how they got left out. Ask for the accuracy at the edges, not the mean. If the vendor has never measured it, that is the answer.
This is why disability belongs in this chapter specifically: it is the cleanest demonstration of the chapter's argument that bias is not a data-cleaning problem. You cannot balance the dataset your way to fairness for disability, because disability is not one group. It is thousands of different, often individually rare, ways of being outside the assumed norm, and the very logic of optimizing for the majority is what produces the harm. Canadian inclusive-design researcher Jutta Treviranus, who directs OCAD University's Inclusive Design Research Centre in Toronto, has made this argument for years: machine learning, by fitting the statistical average, structurally disadvantages the people at the edges of the distribution, which is an apt description of disability. The algorithmic-fairness toolkit from Section Three struggles here because it usually needs a defined group to measure across, and disability resists that definition.
The Canadian frame is real but under-operationalized. The federal Accessible Canada Act (2019) and Ontario's accessibility legislation establish a right to accessibility, and the Canadian Human Rights Act prohibits disability discrimination, but none of them squarely requires that an automated decision system be tested for, or held accountable to, accessibility and non-discrimination by design. The disability movement's founding principle, nothing about us without us, names the missing ingredient: disabled people in the room when these systems are built and procured, not consulted after the harm. It is the same participation gap this chapter keeps finding, in one more place.
This is also, for once, a place where Canada has just moved from naming the gap to writing a rule. In December 2025 Accessibility Standards Canada published CAN-ASC-6.2, billed as the world's first national standard on accessible and equitable AI, written under the Accessible Canada Act by a technical committee made up mostly of people with disabilities. It is unusually concrete about exactly the harms above. It tells organizations to include disabled people at every stage of an AI system, from design to procurement to auditing; to put disabled people in the training data and "share performance results separately for people with disabilities," which is the goblin's accuracy-at-the-edges demand turned into a requirement; and to give people "the choice not to use AI," including a human decision-maker or human review, with the example that a Deaf person must be able to choose a human interpreter over an AI one in a hospital or a court. It draws one bright line straight through the proctoring and hiring problems: organizations should avoid "using AI to judge people based on their body, movement, facial expressions, or emotions."
Here is the catch, and it is the one this book keeps finding. CAN-ASC-6.2 is a standard, not a statute. It is voluntary, a benchmark that could inform a future regulation under the Accessible Canada Act but does not yet bind anyone, and the Act it sits under reaches federally regulated organizations, not the private vendor selling an exam-proctoring tool to a university. Set it beside the one rule Canada does enforce on automated monitoring of workers, Ontario's requirement that employers with twenty-five or more staff keep an electronic-monitoring policy, and the gap is plain: that rule makes employers disclose that they monitor, but grants no right to privacy and sets no limit on what the monitoring data is used for. Strong principles, soft enforcement, which is the whole enforceability ladder in a single disability file.
The neurodivergent edge of this is sharp and under-discussed. When a 2024 study fed near-identical résumés to a leading AI model, the model ranked the versions that mentioned a disability honour lower, and autism fared worst of the disabilities tested, with the system rationalizing that one candidate had shown "less emphasis on leadership." The same logic runs through AI productivity monitoring: tools that score "focus," keystroke rhythm, or on-camera attention are calibrated to a neurotypical baseline, so the focus-then-break working pattern of someone with ADHD, or an autistic worker's different affect, reads to the software not as a different way of working but as a worse one. It is the edge-exclusion Treviranus described, now writing performance reviews.
---
The epistemology layer
This brings the chapter to the Abundant Intelligences paper's central claim, introduced in Chapter 1: "The AI industry-academic complex does not have an ethics problem. It does, however, have an epistemology problem."
Working through what this means in practice, drawing on the Lewis/Whaanga/Yolgörmez analysis:
The foundational definitions of "intelligence" in AI research embed specific historical and cultural assumptions. The dominant working definition, Legg and Hutter's 2007 formulation of intelligence as "an agent's general ability to achieve goals in a wide range of environments," descends from a citation lineage that includes the 1994 "Mainstream Science on Intelligence" statement organized by Linda Gottfredson, which was published as a defence of the IQ-research tradition against critics of The Bell Curve. The citation lineage behind AGI research's working definition of intelligence runs through a document containing explicit racial-bell-curve content — that is the Lewis/Whaanga/Yolgörmez finding, and they quote the document directly. Their argument is not that AGI researchers are personally racist; their argument is that the field's foundational concept of its central topic incorporates assumptions from a particular intellectual tradition without examining where those assumptions came from.
The same pattern operates at the metric layer. The benchmarks AI systems are evaluated against (accuracy on specific tasks, performance on standardized test corpora, similarity to expert human judgment) were assembled by particular communities with particular concerns. They do not neutrally measure "AI capability." They measure the capabilities those particular communities chose to care about. Achievements on the chosen benchmarks tell you the system does well on what the benchmark designers cared about. They tell you less about whether the system does well on what other communities, with other concerns, would care about.
And the same pattern operates at the deployment layer. The decisions about which AI applications to build, which problems to solve with AI, and which users to design for, are made by specific people in specific institutions with specific resources. The fact that healthcare AI systems are dramatically better at diagnosing conditions common in well-funded healthcare systems' patient populations, and less good at diagnosing conditions common in underserved populations, is not a coincidence and not primarily a technical failure. It reflects the upstream choices about which datasets to assemble, which conditions to target, which users to optimize for.
The framing the Lewis/Whaanga/Yolgörmez analysis produces: the "AI ethics problem" framing treats AI as a fundamentally sound technology that occasionally produces problematic outputs requiring correction. The "AI epistemology problem" framing treats the apparent ethics problems as visible consequences of deeper structural choices: about what counts as intelligence, what counts as success, and ultimately whose intelligence counts at all. The ethics framing supports incremental improvement within current foundations; the epistemology framing supports changing those foundations.
These two framings are not interchangeable. They point toward different policy agendas and produce different readings of the same documented harms — and, not incidentally, they tend to bring different people to the table. The guide notes the distinction without resolving which framing is correct. That resolution is contested, and the guide's methodology asks readers to engage the contest rather than collapse it.
And because the guide's method has to apply to the guide's own organizing move: the epistemology framing has serious critics, and they get their paragraph. The strongest counter-readings run roughly as follows. A citation lineage may prove less than it appears to: working AI researchers optimize benchmarks and loss functions, not Gottfredson's definition, so the genealogy may explain the field's blind spots without describing its daily practice. The algorithmic fairness literature has been engaging power and structure more directly than the ethics-versus-epistemology binary gives it credit for. And "change the foundations" remains, so far, a research program rather than a demonstrated alternative at scale; the Abundant Intelligences prototypes are real, and they are also small. None of this defeats the Lewis/Whaanga/Yolgörmez argument. It does mean the argument is a contested position the guide finds persuasive, not a settled fact the guide gets to build on for free. The guide builds on it anyway. Lean declared.
A note on what this framing is and is not. Lewis, Whaanga, and Yolgörmez are not arguing against AI as such. Their paper is explicit that they diverge from the "decomputerization" camp, the position that AI development should stop or be substantially restricted. Their proposal is to engage AI from different epistemological foundations rather than to refuse it. The Abundant Intelligences research program is concrete technical work: Hawaiian language object recognition prototypes, Lakota hardware-building protocols, Indigenous-language NLP, frameworks for AI built on different assumptions about what intelligence is. The epistemology critique points toward constructive alternative practice, not toward abandonment. Treating the critique as anti-AI mischaracterizes it.
---
AI in the classroom: the fairness test Canadians meet first
For most Canadians, the first place AI starts making decisions about them is school, their own or their kid's, and it is a good place to watch the bias question move from theory onto a report card.
Start with governance, because the gap explains the rest. Education in Canada is exclusively provincial, and on generative AI the provinces have mostly not acted. As of 2026 only British Columbia, Quebec, and New Brunswick have issued government guidance for K-12 schools, and all three are explicitly advisory, handing the real decisions to boards and individual teachers. The Council of Ministers of Education, the body whose job is to coordinate nationally, has held one closed discussion and produced nothing. The federal AIDA bill never mentioned education at all. The Canadian Teachers' Federation has spent two years asking someone, anyone, to set rules, naming the risks it sees: surveillance, commercial harvesting of student data, the digital divide, and teacher workload. Into that vacuum, vendors sell. It is worth noting what the structural opposite looks like: in 2026 Norway's government recommended that AI mostly stay out of the early grades, age-tiered, with carve-outs for accessibility. You can disagree with that call, but it is a call; Canada has mostly made none.
The vacuum's sharpest edge is detection, and here the evidence is unusually clean. The tools schools reach for to catch AI-written work do not reliably work, and they fail in a biased direction. OpenAI quietly killed its own AI-text detector in 2023 after it correctly flagged just 26 percent of AI text. A peer-reviewed Stanford study the same year found that popular detectors flagged more than 61 percent of TOEFL essays by non-native English writers as AI-generated, while almost never misfiring on essays by US students, because the detectors read simpler vocabulary and sentence structure as machine-like. That is a textbook case of the structural bias this chapter is about: a tool that punishes a group for how it writes, not for what it did. Canadian universities noticed. UBC declined to turn on Turnitin's AI detector and reaffirmed the decision; Waterloo switched it off in 2025, citing unreliability and bias against students whose first language is not English; Guelph tells instructors not to run student work through detectors at all. The same false-positive mechanism worries accessibility advocates about neurodivergent students, though that specific harm has not yet been measured.
Surveillance is the other edge, and Canada has a marquee case. During the pandemic, universities adopted AI proctoring tools that watch students through their webcams; the facial detection in one was shown to miss Black faces more than half the time, the by-now familiar pattern. When a UBC learning-technology specialist, Ian Linkletter, publicly criticized one such vendor, the company sued him. The case ran five years through the BC courts before ending in a consent dismissal in late 2025 with no money paid, a long chilling shadow over anyone who would scrutinize classroom surveillance software. The student-data layer underneath is governed unevenly: Quebec's Law 25 now requires parental consent for children's data and a human-review right for automated decisions, but a 2025 breach of one school-records vendor exposed records on millions of Canadians, and regulators found boards had not written privacy terms into their contracts.
None of this means the technology is useless in a classroom, and the honest account has to say so. The single most rigorous study to date, a randomized trial at Harvard, found students learned more than twice as much from a carefully built AI tutor than from a strong active-learning class. Read the fine print, though, because it is the whole story: that tutor was heavily engineered by physicists, the study ran for one session on one topic, and it is not the off-the-shelf chatbot a school actually buys. When a real product met real students at scale, the result was more sobering: Sal Khan conceded his widely promoted tutoring bot was, for many students, "a non-event," and the usual sales pitch leans on a 1984 tutoring statistic that later research cut roughly in half. There is, as yet, no Canadian study of comparable rigour either way.
The people who study this for a living have mostly converged on the same answer, and it is not a detector. The University of Calgary's Sarah Eaton, the country's leading academic-integrity scholar, argues that running students' work through a secret detector is itself an integrity breach by the instructor, and that the durable response is to redesign how learning is assessed rather than to escalate an unwinnable detection arms race. It is the move this guide keeps arriving at from other directions: when the tool cannot be trusted and the surveillance carries costs, the question stops being "how do we catch them?" and becomes "what were we actually trying to measure?"
---
The corporate ethics-washing pattern
A specific structural pattern deserves naming directly because it shapes how readers should interpret most published AI ethics material from large corporations.
Major AI companies (OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, IBM, and increasingly Cohere) have published AI ethics frameworks, principles documents, governance structures, and (in some cases) external advisory boards. The frameworks share common features: commitments to fairness, transparency, accountability, safety, privacy, beneficial use, and human oversight. The frameworks read well. The question is what they actually constrain.
A specific working test. When a company's commercial interests align with a stated ethical commitment, the commitment gets implemented strongly. When the commercial interests conflict with the commitment, the commitment gets implemented weakly or selectively. This is not a scandal; it's the predictable behaviour of ethics frameworks that answer, in the end, to commercial constraints. The corporation's primary obligation is to its shareholders or investors, not to its ethics framework. The framework operates as a guide where it doesn't substantially constrain commercial activity, and yields where it does.
Documented examples:
OpenAI's safety commitments and the GPT-4o rollback. OpenAI publishes substantial commitments to safety testing, gradual deployment, and red-team review. In April 2025, OpenAI deployed a GPT-4o update that produced documented concerning behaviour patterns (excessively sycophantic responses, validation of user beliefs without appropriate pushback), and rolled back the update only after public criticism. The internal safety review processes the company's published framework would seem to require did not prevent the deployment. The framework is real; the framework also did not actually prevent the documented harm in this case.
Google's AI ethics advisory board and its 2019 dissolution. Google announced an external AI ethics advisory board in 2019 and dissolved it within a week after public objections to specific board members. The dissolution illustrates a pattern: external ethics governance is added when it provides reputational benefit and removed when it produces controversy that outweighs the reputational benefit. The underlying commitment to external oversight was, in practice, contingent on the absence of controversy that the oversight could produce.
The pattern across Stanford FMTI findings. The Stanford Foundation Model Transparency Index (Chapter 18) documented in its 2025 edition that mean transparency scores across major model developers dropped 17 points from 2024, meaning the companies that publish ethics commitments became measurably less transparent over the year in which AI deployment grew most rapidly. The commitments to transparency stayed on the websites; the transparency in practice declined.
Anthropic's public-benefit corporation structure and investor concentration. Anthropic is structured as a public-benefit corporation with explicit safety-and-alignment commitments. In 2024–2025, large incumbents (Amazon, Google) made major reported investments in the company — raising the question of how that capital structure interacts with the company's governance and decision-making. The structural commitment is real; the practical implications of the underlying capital structure are not fully resolved.
How to read corporate AI ethics frameworks. They are not entirely cynical exercises. Many of the people writing them, working on them, and trying to enforce them are sincere. The frameworks have produced real changes in some contexts. They are also not equivalent to binding regulation, independent oversight, or structural change. Reading corporate ethics frameworks as the primary mechanism for addressing AI's structural problems mistakes commitments-on-paper for real constraints. The guide asks readers to treat corporate ethics frameworks as one input, corporate self-presentation, alongside the independent academic literature, the civil society documentation, the litigation record, and the empirical study of actual deployment outcomes. Where the framework and the operational evidence diverge, the operational evidence is the more reliable signal.
ALIGNMENT — does the principle give anything up? A company's AI principles are easy to write and easy to admire. The test is what they cost: find the moment where the stated commitment would have blocked a profitable move, and check whether it actually did. A principle that always yields when money is on the other side is marketing wearing the costume of governance.
---
The Abundant Intelligences constructive program
The chapter has named the epistemology critique and the corporate ethics-washing pattern. It would be unsatisfying to leave readers there. The Lewis/Whaanga/Yolgörmez research program is doing concrete technical and conceptual work that points toward what alternatives could look like, engaged in practice rather than as decoration.
The Abundant Intelligences program coordinates research pods at multiple institutions. It is co-led from Concordia University (Jason Edward Lewis) and Massey University in Aotearoa New Zealand (Hēmi Whaanga), with a University of Lethbridge pod drawing on Niitsitapi (Blackfoot) knowledge keepers and community members, and further pods at Bard College in New York and the University of Hawai'i at West O'ahu. The program is funded through a New Frontiers of Research Fund Transformation grant of more than $22 million from the federal government, a significant federal commitment to Indigenous-led AI research that runs in parallel to the AI for All strategy without being prominently featured in it.
The research program is organized around three principles drawn from Indigenous epistemological traditions:
Regeneration. Building AI systems oriented toward leaving more for future generations than is taken. This contrasts with the extractive default in current AI: large training data extraction, large compute and material consumption, large electricity draw. Regeneration as a design principle asks: what AI would be built if the requirement was that the system's existence leaves the underlying resources, communities, and relationships better than it found them?
Generosity. Building AI oriented toward sharing with other beings in the location where the system operates: humans, non-human animals, lands, waters, and the broader ecosystem of relationships. Where current AI optimizes for narrow objectives, generosity reframes the design question entirely — what would get built if the system had to serve a community of beings rather than a narrow class of users?
Reciprocity. Building AI oriented toward mutual sustenance — relationships in which the system gives back to what supports it. This contrasts with the unidirectional extraction default in current AI. Reciprocity as a design principle asks: what AI would be built if the requirement was that the system returns value to the people, communities, and resources that made it possible?
These principles are not abstract. The Abundant Intelligences program is producing specific technical artifacts:
Hua Ki'i. A Hawaiian-language object recognition prototype, developed in collaboration with Caroline Running Wolf and Noelani Arista, that includes Hawaiian protocols in its development from cultural-protocol grounding through technical implementation. The prototype is a working demonstration that AI built on different epistemological foundations is technically possible, not just philosophically argued for. Hua Ki'i recognizes objects relevant to Hawaiian cultural and ecological contexts, with object categories and naming conventions developed through community process rather than imposed from external benchmarks.
Indigenous-language NLP. Multiple subprojects within the broader program focus on natural language processing for Indigenous languages, including languages in the Niitsitapi-Haudenosaunee Canadian pod's scope, with training data assembled through community-controlled processes rather than through web scraping. The technical challenges (small-data NLP, low-resource language modelling, model architectures suited to morphologically rich languages) are real and the work is producing both technical contributions and frameworks for thinking about consent-and-governance in Indigenous-language AI.
Suzanne Kite's hardware protocols. As discussed in Chapter 4, Suzanne Kite's "How to Build Anything Ethically" essay in the Indigenous Protocol and AI Position Paper applies Lakota construction protocols to building physical computing hardware. The framework remains a conceptual contribution rather than a fully-implemented hardware production system, but it provides specific operational criteria for what materially ethical hardware-building would involve.
The broader 2020 Indigenous Protocol and AI Position Paper. Beyond the academic Abundant Intelligences paper, the broader Position Paper assembled contributions from multiple Indigenous-led researchers and practitioners across Aotearoa, North America, the Pacific, and beyond. The Paper is deliberately polyphonic (design guidelines, scholarly essays, short fiction, poetry, prototype documentation, art), reflecting the explicit position that Indigenous voices on AI are heterogeneous rather than singular. The methodological commitment to polyphony itself is a substantive contribution.
Why this matters for Canadian AI policy. The federal government is funding Abundant Intelligences at substantial scale (a more-than-$22-million NFRF grant) while AI for All does not prominently engage Indigenous-led AI research. The parallel structures are notable. The government supports the alternative epistemology research program through one funding mechanism while pursuing a strategy that doesn't substantively engage that program's frameworks. Whether the two should be integrated, whether the parallel structure is appropriate, and what an integrated Canadian AI strategy that engaged both would look like, are open policy questions Chapter 20 returns to.
---
The working test
The working test the guide asks readers to apply to any AI ethics, bias, or fairness claim:
Sense. Which sense of "bias" is being used — statistical, social, or structural? Does the discussion conflate them?
Specificity. What specific deployment, dataset, or system is being analyzed? Or is the claim abstract enough to resist evaluation?
Stakeholders. Whose values are encoded in the analysis? The corporate ethics team's? The academic researchers'? The affected community's? Who decided what counts as fair?
Solutions. What kind of solution is being proposed — technical fix, governance change, structural reform, alternative framework? What does each solution actually constrain?
Verification. Are the claims about bias, fairness, or ethics commitments independently verifiable, or are they self-reported by the deploying entity?
Foundations. Does the analysis engage the deeper epistemological questions about whose intelligence counts and whose interests AI serves, or does it work within current foundations?
The test is structurally consistent with the working tests in earlier chapters. The methodology is the same because the underlying analytical work is the same: name what's actually being claimed, name the gaps and the foundational assumptions underneath them, then let readers reason through their own positions.
---
Past 'just clean the data'
CHAPTER RECAP — you now have: - The three distinct senses of "bias" in AI — statistical, social, structural — and the methodological commitment to track which sense is operative in any specific claim. - Documented Canadian cases of algorithmic bias — Clearview AI, Toronto Police risk prediction, federal benefits algorithms, healthcare AI, hiring AI — as the empirical baseline for what the bias conversation is actually about. - The algorithmic fairness literature engaged honestly — real contributions (mathematical impossibility results, formal definitions, bias-measurement tools, real improvements in specific contexts), real limits (output-layer focus, protected-category focus, technical-solution focus that doesn't engage structural questions). - Disability as the clearest case that bias is not a data-cleaning problem — AI as real assistive benefit and as structural harm (atypical speech, faces, and behaviour misread; benefits and hiring systems failing at the edges), with Treviranus's "design for the edges" critique and the Canadian accessibility-law frame now taking its first concrete form in CAN-ASC-6.2 (the world's first accessible-AI standard, still voluntary). - The Lewis/Whaanga/Yolgörmez "epistemology problem, not ethics problem" framing developed operationally — the foundational definitions of intelligence in AI research, the metric layer, the deployment layer all embed assumptions from particular intellectual traditions that the field's ethics conversation generally doesn't examine. - The corporate ethics-washing pattern named with documented examples — OpenAI's GPT-4o rollback, Google's 2019 board dissolution, Stanford FMTI's documented transparency decline, Anthropic's public-benefit corporation structure and investment pressures — establishing that corporate AI ethics frameworks are one input alongside others rather than the primary mechanism for addressing AI's structural problems. - The Abundant Intelligences constructive program — three principles (regeneration, generosity, reciprocity), specific technical artifacts (Hua Ki'i, Indigenous-language NLP, Kite's hardware protocols), the more-than-$22-million NFRF federal commitment to alternative AI research running in parallel to AI for All — as the substantive alternative engaged operationally rather than decoratively. - The working test for evaluating any AI ethics claim: sense, specificity, stakeholders, solutions, verification, foundations.
The next chapter (Chapter 16) takes the structural-bias and political-economy analysis this chapter has developed and goes deep on the specific question of AI and labour — what the academic-economics literature shows, what the union documentation shows, what the federal strategy commits to, and where the genuine disagreements lie. The Acemoglu vs Autor debate enters there as the parallel to this chapter's algorithmic-fairness-literature presentation: rigorous academic disagreement engaged honestly rather than collapsed.
You can now read any AI ethics, bias, or fairness claim with the structural equipment to recognize which layer is being engaged, what's being left unexamined, and where the deeper questions sit. The conversation about AI ethics in Canadian public discourse oscillates between corporate-ethics-statement framing (where the bias problems are real but the proposed solutions don't touch the structural drivers) and dismissive framing (where the bias concerns are treated as exaggerated or as already-solved). The empirical evidence and the substantive intellectual work both sit in the middle, with the Abundant Intelligences program demonstrating that alternative foundations are operationally possible rather than merely theoretically argued for.
---
Bias label for this chapter: structural-political analysis of AI ethics with explicit commitment to the epistemology-not-ethics framing as the manual's organizing intellectual move. Author lean: skeptical of corporate ethics frameworks as the primary mechanism for addressing AI structural problems; sympathetic to Indigenous-led alternative epistemological frameworks as substantive contributions rather than as diversity-checkboxes; willing to name the algorithmic fairness literature as doing real work within methodological limits that the literature itself sometimes encourages; explicit that the manual is not Indigenous-led and engages Indigenous-led frameworks (Abundant Intelligences, Kite, the broader Position Paper) through respectful translation rather than authoritative summary. Corporate self-disclosure (OpenAI, Google, Anthropic, Meta) labelled and read accordingly. Academic peer-reviewed sources (Buolamwini and Gebru, Obermeyer et al., Kleinberg/Mullainathan/Raghavan, Lewis/Whaanga/Yolgörmez, the broader algorithmic fairness literature) treated as primary on their respective contributions. Indigenous-led primary sources (the Abundant Intelligences paper, the Indigenous Protocol and AI Position Paper) presented in their own terms with operational implications named.
Primary sources cited or relied on in this chapter: Lewis, Whaanga & Yolgörmez, "Abundant intelligences: placing AI within Indigenous knowledge frameworks," AI & Society 40(1):2141–2157, 2024; Indigenous Protocol and AI Position Paper (2020); Buolamwini & Gebru, "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification" (Proceedings of Machine Learning Research, 2018); Obermeyer et al., "Dissecting racial bias in an algorithm used to manage the health of populations" (Science, 2019); Kleinberg, Mullainathan & Raghavan, "Inherent Trade-Offs in the Fair Determination of Risk Scores" (ITCS 2017); Chouldechova, "Fair prediction with disparate impact" (FAccT 2017); NIST Face Recognition Vendor Test (FRVT) ongoing series; Office of the Privacy Commissioner of Canada joint findings on Clearview AI (2021); Office of the Privacy Commissioner of Canada 2024-25 privacy opinion research documenting 88% concern about AI training data use and 28% Big Tech trust; Stanford Center for Research on Foundation Models, 2025 Foundation Model Transparency Index; New Frontiers of Research Fund Transformation grant documentation on the Abundant Intelligences program (more than $22 million); Treasury Board of Canada Secretariat AI Register documentation; AI for All strategy documentation (June 2026). Detailed citations in the Sources appendix.
---
🧌 GOBLIN CHECK — A system that's 99% accurate on average and 90% accurate on your face is not, for you, 99% accurate. Averages are where disparities go to do public relations. Whenever someone quotes one accuracy number for a system used on millions of different faces, ask for the breakdown. Buolamwini and Gebru asked for the breakdown, and the breakdown was the whole story.
🧌 GOBLIN CHECK — When a vendor says a system is "99% accurate," the goblin's first question is for whom? Accuracy is an average, and disabled people are exactly the bodies an average is built by setting aside. "Works for most people" is no reassurance to the person it doesn't work for. It is a description of how they got left out. Ask for the accuracy at the edges, not the mean. If the vendor has never measured it, that is the answer.
Recap
- The three distinct senses of "bias" in AI — statistical, social, structural — and the methodological commitment to track which sense is operative in any specific claim.
- Documented Canadian cases of algorithmic bias — Clearview AI, Toronto Police risk prediction, federal benefits algorithms, healthcare AI, hiring AI — as the empirical baseline for what the bias conversation is actually about.
- The algorithmic fairness literature engaged honestly — real contributions (mathematical impossibility results, formal definitions, bias-measurement tools, real improvements in specific contexts), real limits (output-layer focus, protected-category focus, technical-solution focus that doesn't engage structural questions).
- Disability as the clearest case that bias is not a data-cleaning problem — AI as real assistive benefit and as structural harm (atypical speech, faces, and behaviour misread; benefits and hiring systems failing at the edges), with Treviranus's "design for the edges" critique and the Canadian accessibility-law frame now taking its first concrete form in CAN-ASC-6.2 (the world's first accessible-AI standard, still voluntary).
- The Lewis/Whaanga/Yolgörmez "epistemology problem, not ethics problem" framing developed operationally — the foundational definitions of intelligence in AI research, the metric layer, the deployment layer all embed assumptions from particular intellectual traditions that the field's ethics conversation generally doesn't examine.
- The corporate ethics-washing pattern named with documented examples — OpenAI's GPT-4o rollback, Google's 2019 board dissolution, Stanford FMTI's documented transparency decline, Anthropic's public-benefit corporation structure and investment pressures — establishing that corporate AI ethics frameworks are one input alongside others rather than the primary mechanism for addressing AI's structural problems.
- The Abundant Intelligences constructive program — three principles (regeneration, generosity, reciprocity), specific technical artifacts (Hua Ki'i, Indigenous-language NLP, Kite's hardware protocols), the more-than-$22-million NFRF federal commitment to alternative AI research running in parallel to AI for All — as the substantive alternative engaged operationally rather than decoratively.
- The working test for evaluating any AI ethics claim: sense, specificity, stakeholders, solutions, verification, foundations.
Sources
- Lewis, Whaanga & Yolgörmez, "Abundant intelligences: placing AI within Indigenous knowledge frameworks," AI & Society 40(1):2141–2157, 2024
- Indigenous Protocol and AI Position Paper (2020)
- Buolamwini & Gebru, "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification" (Proceedings of Machine Learning Research, 2018)
- Obermeyer et al., "Dissecting racial bias in an algorithm used to manage the health of populations" (Science, 2019)
- Kleinberg, Mullainathan & Raghavan, "Inherent Trade-Offs in the Fair Determination of Risk Scores" (ITCS 2017)
- Chouldechova, "Fair prediction with disparate impact" (FAccT 2017)
- NIST Face Recognition Vendor Test (FRVT) ongoing series
- Office of the Privacy Commissioner of Canada joint findings on Clearview AI (2021)
- Office of the Privacy Commissioner of Canada 2024-25 privacy opinion research documenting 88% concern about AI training data use and 28% Big Tech trust
- Stanford Center for Research on Foundation Models, 2025 Foundation Model Transparency Index
- New Frontiers of Research Fund Transformation grant documentation on the Abundant Intelligences program (more than $22 million)
- Treasury Board of Canada Secretariat AI Register documentation
- AI for All strategy documentation (June 2026).