AI in breast cancer detection is software that reads screening mammograms, scores each exam for signs of cancer and marks suspicious areas. Screening programs use it as a second reader, to decide which exams need a second radiologist, or as a safety net, while radiologists still make the recall decision.

The evidence has grown quickly. In Sweden's randomized MASAI trial, AI-supported screening of 105,934 women found 29% more cancers with 44% fewer screen readings and no significant rise in false positives. This guide explains how the software works, what MASAI, Germany's PRAIM study and the UK's EDITH trial show, which products the US Food and Drug Administration (FDA) lists, where AI is reaching pathology and treatment planning, and where it falls short. It is health information, not medical advice: questions about your own screening belong with your doctor or screening program. AI for lung CT, drug discovery and consumer health apps are separate subjects and get only this mention here.

How AI reads a mammogram

A standard screening mammogram is four X-ray images, two views of each breast. Some services also use digital breast tomosynthesis (DBT), which takes X-rays from several angles and builds a stack of thin images, so overlapping tissue hides less.

The tools in the studies below are deep learning models. PRAIM's system, for example, combines deep convolutional neural networks trained on mammograms labelled normal, benign or malignant. For each exam, a typical product returns three things:

Software has read mammograms before. The FDA approved computer-aided detection (CAD) for mammography in 1998, and by 2015 CAD was used on most US screening mammograms. That year a study in JAMA Internal Medicine compared 495,818 mammograms read with CAD and 129,807 without, and found CAD improved accuracy on no measure: sensitivity was 85.3% with it and 87.3% without. That history is why the newer tools were tested inside real screening programs, on the outcomes that matter: cancers found, women recalled, and cancers missed.

If you have trained a classifier on the Wisconsin breast cancer dataset in a Python tutorial, that is a teaching problem built from measurements of biopsy cells. Screening AI reads the images themselves.

Where AI fits in a screening read

In Sweden, Germany and the UK, screening mammograms are normally read by two radiologists independently, and exams either of them flags go to a consensus discussion or arbitration. The UK government describes it plainly: two specialists are needed per mammogram. AI changes who reads what, in one of four patterns.

PatternWhat the AI doesWho decidesStudied in
Triage by risk scoreSends low-scoring exams to one radiologist and the highest-scoring to twoRadiologistsMASAI, Sweden
Replacing the second readerReads independently in place of one radiologistRadiologists at consensusScreenTrustCAD, Sweden; EDITH, UK
Normal triage and safety netTags confident normals, and alerts a reader who calls a suspicious exam normalThe reading radiologistPRAIM, Germany
Concurrent reading aidShows scores and marks while the physician readsThe interpreting physicianUS labelling of cleared tools

In MASAI, Transpara scored every exam. Exams scoring 1 to 9 went to a single radiologist and exams scoring 10 went to two. Readers saw the risk score on every exam, and the AI's marks on exams scoring 8 to 10. Most exams therefore needed one read instead of two, which is where the trial's workload saving came from.

Screening exams pass through an AI gear that sends most of them to a single reading monitor and one orange, high-scoring exam to a desk with two monitors for double reading.
Fig. 1 Triage moves radiologist time to the few exams most likely to hold a cancer, instead of reading every exam twice.

PRAIM used Vara's system, a CE-certified medical device, differently. The software tagged exams it judged highly unsuspicious as normal in the worklist. For exams it judged highly suspicious, it stayed silent until the radiologist had read the exam; if the radiologist called it normal, a safety-net alert showed the suspicious region and asked for a second look, which the radiologist could accept or reject. In the AI group, the safety net fired on 3,959 exams and was accepted 1,077 times, and 204 breast cancers were diagnosed among those reassessments.

A reading monitor shows an exam already marked normal with a check mark; an AI gear then places an orange ring on one image tile, and a curved arrow sends the exam back for a second look.
Fig. 2 A safety net speaks only after the radiologist has decided, so the first read stays independent of the AI.

In the United States, cleared tools are labelled as aids to the physician. Transpara's clearance, for example, describes a concurrent reading aid for physicians interpreting screening mammograms and DBT, and states that patient management decisions should not rest on its analysis alone.

What the MASAI, PRAIM and other trials found

MASAI: the first randomized trial

MASAI randomly assigned women attending screening at four sites in southwest Sweden, between April 2021 and December 2022, to AI-supported screening or standard double reading. It has reported three times:

  1. The Lancet Oncology, August 2023. The planned safety analysis of the first 80,033 women: 244 cancers found with AI against 203 without (6.1 versus 5.1 per 1,000), a recall rate of 2.2% against 2.0%, a false-positive rate of 1.5% in both groups, and 44.3% fewer screen readings.
  2. The Lancet Digital Health, February 2025. With all 105,934 women: 6.4 versus 5.0 cancers per 1,000, a 29% increase. The extra invasive cancers were mainly small and lymph-node negative. In situ cancers rose from 45 to 68, about half of the increase high-grade, with no increase in the lowest grade. Recall and false-positive rates were not significantly higher, and the reading workload fell by 44.2%.
  3. The Lancet, January 2026. The primary endpoint, interval cancers after two years: 1.55 per 1,000 with AI against 1.76 without, which met the trial's non-inferiority test. Sensitivity was 80.5% against 73.8%, consistent across age and breast density, and specificity was 98.5% in both groups.

The interval cancer difference was not statistically significant: MASAI was designed to show that AI-supported screening is not worse than two radiologists, and it showed that, while finding more cancers with far less reading.

PRAIM: real-world use across Germany

PRAIM, published in Nature Medicine in January 2025, followed 463,094 women aged 50 to 69 screened at 12 German sites by 119 radiologists between July 2021 and February 2023, of whom 260,739 were read with AI support. Radiologists using AI found 6.7 cancers per 1,000, 17.6% more than the 5.7 per 1,000 without it, while the recall rate was 37.4 per 1,000 against 38.3. The share of recalls that turned out to be cancer rose from 14.9% to 17.9%.

PRAIM was observational: radiologists chose, exam by exam, whether to use the AI viewer. The authors name a "reading behavior bias", in which some radiologists preferred the AI viewer for exams tagged normal, and adjusted their analysis for it. The study was funded by Vara, which took part in the design and the writing.

ScreenTrustCAD, GEMINI and the NHS studies

Study, journalWhere and designWomenMain result
MASAI, The Lancet (2026)Sweden, randomized105,93429% more cancers, 44% fewer reads, non-inferior interval cancers
PRAIM, Nature Medicine (2025)Germany, observational463,09417.6% higher detection, no rise in recalls
ScreenTrustCAD, Lancet Digital Health (2023)Stockholm, paired readers55,581AI plus one radiologist non-inferior to two, with 4% more cancers
GEMINI, Nature Cancer (2026)Scotland, prospective10,88910.4% more cancers in the main workflow, up to 31% less reading

ScreenTrustCAD compared combinations of two radiologists and the AI reading the same exams: replacing one radiologist with AI was non-inferior and found 4% more cancers. GEMINI, in NHS Grampian, used Kheiron's Mia live and in simulation to model 17 ways of using AI; when the AI recommended recall and double reading did not, extra human review found 11 more cancers.

In the UK, a study of Google's system on 115,973 mammograms from five NHS services found higher sensitivity than the first reader (54.1% against 43.7%) with non-inferior specificity, and AI detected 25.0% of the interval cancers. When the tool was then run live at 12 sites, without affecting care, a distribution shift meant its thresholds had to be recalibrated. That is the lesson retrospective studies cannot teach. A 2020 Nature paper, for example, reported absolute reductions in false positives of 5.7% on US data and 1.2% on UK data; the prospective studies above are the evidence screening programs weigh.

Note

Several of these studies were funded or co-authored by the AI vendor, as the papers declare: PRAIM by Vara, GEMINI with Kheiron staff, the NHS studies with Google, and ScreenTrustCAD with Lunit among its funders. MASAI was funded by the Swedish Cancer Society, Sweden's regional cancer centres and government clinical research funding.

The UK EDITH trial and its status in 2026

The UK government announced EDITH (Early Detection using Information Technology in Health) on February 4, 2025. Nearly 700,000 women booked for routine NHS screening are to be invited at 30 sites, with £11 million from the National Institute for Health and Care Research (NIHR). The question it tests is the reader-replacement pattern: whether one specialist with AI can complete the screening read safely, where two are needed today.

England's National Cancer Plan, in its version updated on September 15, 2026, says EDITH began recruitment in April 2025 and that the government will consider its results. We found no published EDITH results as of September 2026.

Which AI mammography tools the FDA lists

The FDA publishes a list of AI-enabled medical devices it has authorized. The version current as of September 4, 2026 lists decisions up to June 29, 2026 and includes these breast screening tools, among others:

Product (company)What it doesFirst on the listLatest entry
Transpara (ScreenPoint Medical)Detection aid, 2D and DBTNov 21, 2018Nov 25, 2024 (version 2.1.0)
ProFound AI (iCAD)Detection aidOct 4, 2019Nov 8, 2024 (ProFound Detection V4.0)
MammoScreen (Therapixel)Detection aidMar 25, 2020Jun 26, 2026
Genius AI Detection (Hologic)Detection aidNov 18, 2020Jul 31, 2025 (version 2.0)
Lunit INSIGHT MMG (Lunit)Detection aid, 2DNov 17, 2021Apr 23, 2026
Lunit INSIGHT DBT (Lunit)Detection aid, tomosynthesisNov 6, 2023Mar 26, 2026
Saige-Dx (DeepHealth)Detection aidMay 12, 2022Jun 15, 2026
Allix5 (Clairity)Five-year breast cancer risk predictionMay 30, 2025De Novo authorization

The detection aids in the table were cleared through 510(k) and share one FDA product code, for radiological computer-assisted detection and diagnosis software (21 CFR 892.2090, Class II). Allix5 went through De Novo and created a new category, software that predicts future breast cancer risk from a mammogram. Its decision summary is a good example of what labels contain: it is not intended to detect cancer or guide the reading, its output is meant to be considered after the radiologist's interpretation, and scores above 5% should be read with caution because they may significantly overestimate risk, especially above 10%.

The same list reaches beyond mammography: QuantX, a diagnosis aid for breast MRI, was authorized in July 2017, and Koios DS for Breast, which classifies lesions on breast ultrasound, was cleared in July 2019.

Important

Clearance means a device may be marketed for its labelled use. It is not a ranking, and this post does not recommend any product: how a tool performs depends on the scanners, the reading workflow and the women screened, which is why the limits below matter.

AI in breast pathology and treatment planning

After a cancer is found, AI is moving into two later steps: reading the surgical tissue, and planning radiotherapy. In both, the labels leave the decision with the physician or the planner.

In pathology, the FDA cleared ArteraAI Breast on May 4, 2026. It analyzes a scanned H&E-stained slide from the surgically removed tumour, together with age, tumour size and nodal status, and estimates the 5- and 10-year risk of distant metastasis for adults with early-stage HR+/HER2- invasive breast cancer who are candidates for further (adjuvant) treatment. Its label says it assists physicians with prognostic risk-based decisions alongside other clinical and pathological factors.

In radiotherapy, one planning step is contouring: outlining the target and the organs to spare on the planning CT. AutoContour, cleared on March 19, 2026, uses deep learning models that include left and right breast structures, and lets the planner review and edit every contour before it goes to the treatment planning system.

Health systems are funding the groundwork. England's National Cancer Plan commits £96 million to automate histopathology and plans a faster roll-out of AI-assisted pathology, alongside a larger investment in digital diagnostics.

The limits: false positives, breast density and different populations

False positives and overdiagnosis

In MASAI and PRAIM, AI raised detection without clearly lowering false alarms: MASAI's false-positive rate stayed at 1.5% with and without AI, and PRAIM's recall rate was essentially unchanged (37.4 versus 38.3 per 1,000). Fewer recalls show up so far mainly in retrospective data and in simulated workflows, such as some of the variants GEMINI modelled. More detection also raises the question of overdiagnosis: MASAI's extra in situ cancers were about half high-grade, with no increase in the lowest grade, but none of these studies measured deaths, and more cancers found is not the same as more lives saved.

Breast density

The FDA tells patients that dense tissue makes cancers harder to find on a mammogram and raises the risk of breast cancer. Since September 2024, US mammography facilities must give each patient a written summary of their breast density, and reports to providers use four density categories.

Two image tiles side by side: on the left a small bright spot stands out against dark tissue; on the right the same spot, in orange, is half hidden among pale cloud shapes of dense tissue.
Fig. 3 Dense tissue can hide a tumour from human readers and models alike, so check how any tool performs at each density.

The trials are encouraging here: MASAI's gain in sensitivity was consistent across density groups, and PRAIM reports detection that was non-inferior or better across densities. But a single tool can still behave differently. A 2024 study in Radiology ran an FDA-cleared algorithm on 4,855 women with negative DBT exams and found its false-positive risk scores more likely in extremely dense breasts (odds ratio 2.8) than in fatty breasts.

Different populations, scanners and drift

The same Radiology study found false-positive case scores more likely for Black women (odds ratio 1.5) and women aged 71 to 80 (1.9), and less likely for Asian women (0.7), compared with White women and women aged 51 to 60. MASAI did not collect race or ethnicity, so it cannot answer the question for its own population. The UK study of Google's system found no systematic demographic disparities retrospectively, then met a distribution shift when deployed live. The general problem of error rates that differ by group is covered in our guide to fairness in AI decision-making.

The practical reading: a tool's published numbers describe the population, scanners and settings it was tested on. A new service has to measure it again on its own exams, and keep measuring after every software or scanner update.

Who is accountable for an AI-assisted reading

In every deployment above, a qualified person makes the call. Transpara's FDA label says patient management decisions should not be based on it alone, Allix5's says it must not be the sole determinant of clinical decisions, and PRAIM's design left the final recall decision to the radiologists.

Keeping a person in the loop does not make the person independent of the machine. In a 2023 Radiology experiment, 27 radiologists read mammograms with a purported AI system that gave wrong suggestions on some exams. Very experienced readers rated 82.3% of exams correctly when the AI was right, and 45.5% when it was wrong; inexperienced readers fell from 79.7% to 19.8%. The reverse also happens: in an NHS reader study, human arbitration overruled the AI's false recalls, which improved specificity, but also overruled it on some cancers that later surfaced as interval or next-round cancers.

For a screening service weighing AI, the evidence points to a short list of questions to settle first:

  1. Which pattern, triage, second reader, safety net or concurrent aid, and who decides at each step.
  2. Local validation on your own exams, scanners and population, including performance by density and age.
  3. Thresholds set for your recall and workload targets, and a plan to recalibrate them.
  4. Monitoring of detection, recall and interval cancer rates after go-live and after every update.
  5. A record of each AI score, each alert and whether the reader accepted it, so disagreements can be audited.

Outside medicine, the same rules apply to any AI that informs a decision. That is how our AI and automation service designs workflows and assistants: the model proposes, a rule or a named person approves anything that changes data, and every model or prompt change is tested on the same real cases before it ships.