AI in Business, Education & Society
AI is transforming business" and "AI is transforming education" are two of the least useful sentences in circulation, because they are compatible with almost any outcome — a genuine productivity gain, a stalled pilot project, or active harm — and are usually deployed to support whichever conclusion the speaker already wanted. The more useful question is never "does AI have an impact here?" but "what specifically changed, measured how, over what period, compared to what baseline?" This page works through business, education, and society in turn using exactly that discipline: a real, sourced claim, its actual evidential strength, and what that strength does and doesn't license you to conclude.
AI in Business: Adoption Is Not the Same as Value
Two large, recent survey efforts give a consistent, and consistently more sobering, picture than the "AI is everywhere now" headline suggests. Stanford HAI's 2026 AI Index reports that 70% of organisations now use generative AI in at least one business function [1] — genuine, rapid, measured adoption. McKinsey's 2025 State of AI survey finds a similar figure (around two-thirds of organisations regularly using generative AI in at least one function, up from roughly one-third in 2023) [2], with the most common functions being marketing/sales, product/service development, and IT [2]. Both surveys agree on where the story gets less flattering: McKinsey's own respondents report that only a small minority of organisations — around one in twenty — are seeing measurable financial returns from their AI investment, and only about a third have moved a use case past the pilot stage into anything at scale [2]. Adoption, in other words, is real and fast; value realisation is a separate, much slower-moving variable, and a statistic about the first tells you almost nothing about the second. A claim like "88% of businesses now use AI" is technically defensible and almost useless on its own — the operative question for any real decision is always the value question, not the adoption question, and the two get conflated constantly in press coverage of exactly these reports.
AI in Education: A Genuine Opportunity With a Documented Cost
The opportunity side of AI in education is real and worth naming plainly: automated, always-available feedback at a scale no teaching team can match by hand; draft-level accessibility support (translation, reading-level adaptation, text-to-speech) available instantly rather than after a support request; and a tutoring-style interaction that can be repeated as many times as a student needs without cost or embarrassment, closing at least part of the gap between students who can afford a private tutor and those who cannot.
The cost side is equally real, and this site already has a full page working through the evidence for it in depth: Avoiding the AI Short-Circuit covers the strongest current panel evidence (a 30-month study of over 26,000 secondary students) that using generative AI to complete homework raises assignment scores while lowering unassisted exam performance, and sets out the cognitive-science mechanism — desirable difficulty, the testing effect — that explains why an intervention can improve one proxy for learning while actively undermining the thing the proxy was meant to stand in for. Rather than re-deriving that argument here, the point worth making at this level is the general one it's an instance of: an education technology's marketed benefit and its measured effect on the outcome that actually matters (durable understanding, not a submitted grade) are logically independent claims, and evidence for one is never automatically evidence for the other. Treat any confident claim about an AI tool's educational benefit the same way that page treats the "just use AI to help you study" claim — ask what was actually measured, over what timeframe, and whether the measurement was of performance with the tool present or of retained understanding without it.
AI in Society: Bias, Misinformation, and Labour
Beyond any single organisation's use case, three society-level effects have the strongest evidence behind them, and each illustrates a different way an AI system's failure mode can be invisible until someone deliberately goes looking for it.
Algorithmic bias. The clearest documented case remains ProPublica's 2016 investigation of COMPAS, a risk-assessment algorithm used across US courts to score a defendant's likelihood of reoffending. The investigation found that COMPAS was substantially more likely to falsely flag Black defendants as high-risk than white defendants who did not go on to reoffend — 45% of Black defendants were incorrectly labelled high-risk, against 23% of white defendants — despite the algorithm not using race as an input feature at all [3]. That last detail is the important methodological lesson: removing a protected characteristic from a model's inputs does not remove bias from its outputs if other features are correlated with that characteristic, and a system can be biased in its real-world effect while being "fair" by the narrow standard of not directly consulting the attribute in question. This is not a resolved, historical case study — the same structural failure mode (proxy correlation reintroducing exactly the bias a system was designed to exclude) recurs in every subsequent domain an equivalent model gets deployed into, from hiring screens to credit scoring.
Misinformation and synthetic media. The World Economic Forum's Global Risks Report 2024, drawing on a survey of nearly 1,500 global risk experts, ranked misinformation and disinformation — explicitly including AI-generated synthetic content — as the single highest-ranked global risk over the following two years, ahead of extreme weather and armed conflict [4]. The mechanism worth understanding, not just the headline ranking: what has changed is not that misinformation is new, but that the marginal cost of producing convincing, personalised, high-volume synthetic content has collapsed, at the same moment several billion people worldwide were due to vote in national elections [4]. A risk ranking produced by expert survey is itself a claim with a specific evidential status — it tells you what a particular, large panel of risk professionals believes is likely and severe, not a direct measurement of harm that has already occurred — and it's worth noticing that distinction rather than treating "ranked #1 by the WEF" as equivalent to a measured outcome.
Labour-market effects. Eloundou, Manning, Mishkin, and Rock's widely cited analysis estimated that around 80% of the US workforce could have at least 10% of their work tasks affected by large language models, and around 19% of workers could see at least half of their tasks affected [5]. The paper is worth reading carefully rather than summarised into a single scary number, because its own methodology is explicit about what "affected" means: task exposure, not job replacement — a task being faster or different with LLM assistance, which is a much weaker and more defensible claim than "these jobs will disappear," and one the authors are careful to distinguish in their own discussion [5]. Confusing task-level exposure with job-level displacement is one of the most common ways this specific piece of research gets over-claimed in secondary reporting.
Weighing a Claim Against Its Evidence
Notice the pattern across all three sections above: in every case, the useful move was not accepting or rejecting the headline claim, but asking what kind of evidence sits behind it — a survey of self-reported adoption (business), a 30-month panel study (education), a specific empirical audit of one deployed system (COMPAS), an expert-perception ranking (WEF), and a task-exposure estimate explicitly flagged as not the same as a displacement estimate (labour market). None of these evidence types are interchangeable, and none of them license conclusions the others would. A press release citing "88% AI adoption" and a peer-reviewed panel study finding a 20% exam-score penalty are both real, both about AI, and answer completely different questions — treating either as a stand-in for the other is the single most common error in casual discussion of AI's societal impact, and the fastest way to spot it is to ask, before accepting any claim: what, exactly, was measured, and does that measurement actually support the conclusion being drawn from it?
References
- Stanford Institute for Human-Centered Artificial Intelligence (HAI). (2026). The 2026 AI Index Report — Economy chapter. https://hai.stanford.edu/ai-index/2026-ai-index-report/economy
- McKinsey & Company. (2025). The State of AI: How Organizations Are Rewiring to Capture Value. QuantumBlack, AI by McKinsey.
- Angwin, J., Larson, J., Mattu, S., & Kirchner, L. (2016). Machine Bias. ProPublica. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing
- World Economic Forum. (2024). Global Risks Report 2024. https://www.weforum.org/publications/global-risks-report-2024/
- Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2023). GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. arXiv:2303.10130. https://arxiv.org/abs/2303.10130