Quick answer
AI in a polysomnography lab cuts down manual scoring: it splits the night into 30-second epochs, assigns each a sleep stage (W, N1, N2, N3, REM), and detects apneas, hypopneas, desaturations and arousals. The best systems agree with experts roughly as often as experts agree with each other. They don't make a diagnosis, though. Software cleared in the US, such as SleepStageML and SOMNUM v3.0, works under physician supervision, in adults only, and only within the scope it was validated for. A human reviews and signs off the result. We treat the analysis in our pipeline for the MT5 Foundation the same way: as illustrative and research material, not a physician's report.
What gets scored in a PSG study
Polysomnography records the whole night across more than a dozen synchronised channels: EEG, eye movements (EOG), chin muscle tone (EMG), ECG, airflow, chest and abdominal effort belts, oxygen saturation (SpO2), body position and limb movements. Scoring follows the AASM manual (AASM Manual for the Scoring of Sleep and Associated Events). The current version 3 came out in February 2023, and AASM-accredited facilities had to adopt it by the end of 2023.
The main things that get scored:
- Sleep stages. Each 30-second epoch gets one label: wake (W), N1, N2, N3 or REM. The sequence of epochs forms the hypnogram.
- Arousals. An abrupt shift in EEG frequency lasting at least 3 seconds, preceded by stable sleep. In REM, a concurrent increase in chin muscle tone is also required.
- Respiratory events. Apnea (airflow stops) and hypopnea (airflow drops). Apneas are classified as obstructive, central or mixed, depending on whether the chest and abdomen are still trying to breathe.
- Desaturations. A drop in SpO2 of at least 3% or 4% from baseline. The AASM's recommended rule scores a hypopnea with a desaturation of at least 3% or an arousal. In version 3 of the manual, the 4% criterion changed from "acceptable" to "optional".
These elements produce the indices that the study report relies on. For automated analysis, one distinction matters: some rules are purely signal-based (a 3% SpO2 drop can be computed deterministically), while others need pattern interpretation (is this epoch N1 already, or still wake?).
The baseline: people disagree too
Before asking how accurate AI is, you have to ask: accurate compared to what? Sleep scoring has no independent ground truth. There is an expert's judgement, and experts differ.
The largest dataset on this comes from the AASM inter-scorer reliability program. Rosenberg and Van Hout (2013) analysed more than 3.2 million decisions by over 2,500 scorers on 1,800 epochs. Mean agreement with the majority score was 82.6%. It was highest for REM, followed by N2 and wake. For N3 it fell to 67.4%, and for N1 to 63.0%.
Respiratory events look similar. In the 2014 analysis, agreement for epochs with any respiratory event was 88.4%, but only 65.4% for hypopneas (κ = 0.57) and 52.4% for central apneas (κ = 0.41).
Bakker et al. (2023) went further and had the same recordings scored by 6, 9 and 12 experts. The share of epochs where every scorer agreed was 46%, 38% and 32% respectively. The authors proposed reporting stage probabilities (a hypnodensity) instead of one label per epoch. That is a useful frame for thinking about AI: a model that says "this epoch is 60% N1, 40% N2" describes the night more honestly than one that pretends to be certain.
What validation studies show
Below are a few papers worth knowing, because they compare an automated system with people on clinical data. This is not a full literature review.
| Study | What was compared | Result | Caveats |
|---|---|---|---|
| Punjabi et al., Sleep 2015 (Somnolyzer) | 97 recordings scored manually at 4 labs and automatically | AHI correlation machine vs humans 0.93, between labs 0.92; good agreement on arousal index, total sleep time, sleep efficiency | largest differences in % of N1, N2 and N3; 2007 AASM criteria |
| Perslev et al., npj Digital Medicine 2021 (U-Sleep) | model trained on recordings of 15,660 people from 16 clinical studies | on data from a clinic unseen in training, as accurate as the best human expert | sleep stages only, no respiratory events |
| Bakker et al., Sleep 2023 | 95 recordings, each scored by 6-12 experts, hypnodensity comparison | ICC between automated and manual stage probabilities 0.91 | also shows how rare full agreement between experts is |
Two conclusions recur across these papers. First, at the level of whole-night indices (sleep time, AHI, arousal index), machine and humans agree well. Second, at the level of individual epochs, errors cluster where humans are uncertain: stage transitions, N1, sleep depth, hypopnea type.
The AASM's 2020 position statement on AI in sleep medicine names automated PSG scoring as the most immediate practical application, but sets conditions: a clearly stated population and purpose for the tool, validation on independent data, and transparency towards the lab that uses it.
What is cleared for clinical use
In the US, automated PSG analysis goes through the FDA 510(k) pathway. Two recent examples:
- SleepStageML (Beacon Biosignals), 510(k) K233438, decided in March 2024. It scores sleep stages from EEG in PSG recordings of adults and is intended to assist a clinician in evaluating sleep. It was the first sleep medicine device cleared with a predetermined change control plan (PCCP), which allows the model to be updated without a new submission, as long as each version passes the agreed tests.
- SOMNUM v3.0 (HoneyNaps), 510(k) K253390, announced in July 2026. It analyses level 1 PSG recordings of adults aged 22 and over: sleep stages, arousals, limb movements, apneas and hypopneas, classifying apneas as obstructive, central or mixed. Its intended use is under physician supervision, and the physician can edit and delete events.
The AASM also runs its own certification for autoscoring software. The 2023 pilot covered adult sleep staging only, and the first certified product was Sleepware G3 with Somnolyzer. In April 2026 the AASM announced a full-PSG program covering stages, respiratory events, arousals and limb movements, assessed on a multi-site set of recordings. It requires FDA clearance or a pending submission.
In the European Union, software that provides information for diagnostic decisions is usually a medical device. Rule 11 of Annex VIII to the MDR puts it in at least class IIa, which means a notified body is involved. An AI system in such a device is a high-risk system under the AI Act. Regulation (EU) 2026/1744, the AI Omnibus, moved the application date for those obligations for AI in regulated products such as medical devices to 2 August 2028. A tool used only in a research project, with no effect on decisions about an individual patient, is in a different position, but that line has to be drawn deliberately.
This is not legal advice. Classifying a specific piece of software is best done with a medical device regulatory specialist.
What AI does well, and where a human is needed
| Task | Automation | Human role |
|---|---|---|
| Desaturations (SpO2 drop ≥ 3% or 4%) | signal rule, repeatable | check sensor artefacts, a slipped oximeter |
| W, N2, REM | usually at the level of inter-scorer agreement | review the hypnogram, especially at transitions |
| N1 and the N2/N3 boundary | most disagreement, for humans too | decide uncertain epochs |
| Arousals | detected from EEG and EMG, sensitive to signal quality | verify, because they affect hypopnea scoring |
| Apnea and hypopnea type | hard, humans agree poorly | clinical judgement in the patient's context |
| Missing or noisy channels | many models lose quality | decide whether the recording can be scored |
| Populations outside validation (children, neurological disease) | no guarantees | out of the tool's scope, manual scoring |
| Study report and diagnosis | outside the automation's scope | physician only |
In practice, a sensible split looks like this: the software does a first pass over the whole night and flags low-confidence epochs and events, and a technologist or physician reviews those first. The time saving comes from not scoring eight hours of signal from scratch, not from nobody looking at the signal anymore.
How we do it: the pipeline for the MT5 Foundation
The MT5 Foundation runs sleep labs in hospitals. Today they record about a hundred nights a month. We delivered the technical side: hardware, deployment, the data route and an NVIDIA Blackwell GPU cluster that sits at the foundation. The full story is in our MT5 Foundation case study.
The pipeline works like this:
- Recording. A NOX A1s recorder writes the night to a single EDF+ file. The sample night in our story has 90 channels in the file, 8 hours of recording and about 495 million samples.
- Preparation. Anonymisation, filtering (0.3-35 Hz for EEG), a common sampling grid and channel alignment down to the sample.
- Describing the night. The pipeline automatically detects airflow cessations, desaturations of at least 3% and arousals, and derives an illustrative hypnogram from the EEG, EOG and activity. We describe events neutrally, without classifying their type.
- Windows for the model. The night is cut into 30-second windows on which we train our own model. In each window we mask one channel, such as saturation, and the model has to rebuild it from the rest. To do that, it has to learn how breathing, heart, brain and movement depend on one another.
All of this analysis is illustrative and for research. It is not a physician's report and is not used to diagnose patients. The research questions we are working on are open. We have no results yet and will not announce any before rigorous verification. Patient data never leaves the foundation, because the pipeline, the model and the data sit on hardware in its own infrastructure. Our approach to projects like this is described on the R&D page.
Checklist before rolling out automated scoring
- Purpose. Is the tool meant to support clinical reporting, or only research? That determines which clearances you need.
- Population. Does the validation cover your patients: age, comorbidities, study type (in-lab PSG or home testing)?
- Channels and hardware. Was the model validated on your montage and recorder? What does it do when a channel is missing or noisy?
- Rule version. Which version of the AASM manual and which hypopnea criterion (3% or arousal, or 4%)? It has to match what you report.
- Your own validation. Compare the software with at least two human scorers on a few dozen of your recordings. Measure machine-human agreement and human-human agreement.
- Uncertainty. Does the system show which epochs and events are uncertain? Review should start there.
- Audit trail. Can you see what the software did and what a human changed? This matters for quality control and model updates.
- Data. Where are recordings processed? If off site, on what legal basis and with what safeguards? If on site, who maintains the server?
Sources
- AASM: AASM releases updated version of scoring manual (version 3, 2023)
- AASM: Summary of Updates in Version 3 (PDF)
- AASM: AASM clarifies hypopnea scoring criteria
- Rosenberg, Van Hout: The AASM Inter-scorer Reliability Program, Sleep Stage Scoring, JCSM 2013
- Rosenberg, Van Hout: The AASM Inter-scorer Reliability Program, Respiratory Events, JCSM 2014
- Punjabi et al.: Computer-Assisted Automated Scoring of Polysomnograms Using the Somnolyzer System, Sleep 2015
- Perslev et al.: U-Sleep, resilient high-frequency sleep staging, npj Digital Medicine 2021
- Bakker et al.: Scoring sleep with artificial intelligence enables quantification of sleep stage ambiguity, Sleep 2023
- Goldstein et al.: Artificial intelligence in sleep medicine, an AASM position statement, JCSM 2020
- AASM: Beacon Biosignals receives FDA clearance for sleep staging software
- AASM: FDA clears HoneyNaps SOMNUM v3.0 sleep analysis software
- AASM: autoscoring certification pilot program (2023)
- AASM: Full PSG Autoscoring Certification Program (2026)
- MDCG 2019-11: Guidance on qualification and classification of software (MDR/IVDR)
- Regulation (EU) 2026/1744 (AI Omnibus), EUR-Lex
