Quick answer

AI in a polysomnography lab cuts down manual scoring: it splits the night into 30-second epochs, assigns each a sleep stage (W, N1, N2, N3, REM), and detects apneas, hypopneas, desaturations and arousals. The best systems agree with experts roughly as often as experts agree with each other. They don't make a diagnosis, though. Software cleared in the US, such as SleepStageML and SOMNUM v3.0, works under physician supervision, in adults only, and only within the scope it was validated for. A human reviews and signs off the result. We treat the analysis in our pipeline for the MT5 Foundation the same way: as illustrative and research material, not a physician's report.

What gets scored in a PSG study

Polysomnography records the whole night across more than a dozen synchronised channels: EEG, eye movements (EOG), chin muscle tone (EMG), ECG, airflow, chest and abdominal effort belts, oxygen saturation (SpO2), body position and limb movements. Scoring follows the AASM manual (AASM Manual for the Scoring of Sleep and Associated Events). The current version 3 came out in February 2023, and AASM-accredited facilities had to adopt it by the end of 2023.

The main things that get scored:

  • Sleep stages. Each 30-second epoch gets one label: wake (W), N1, N2, N3 or REM. The sequence of epochs forms the hypnogram.
  • Arousals. An abrupt shift in EEG frequency lasting at least 3 seconds, preceded by stable sleep. In REM, a concurrent increase in chin muscle tone is also required.
  • Respiratory events. Apnea (airflow stops) and hypopnea (airflow drops). Apneas are classified as obstructive, central or mixed, depending on whether the chest and abdomen are still trying to breathe.
  • Desaturations. A drop in SpO2 of at least 3% or 4% from baseline. The AASM's recommended rule scores a hypopnea with a desaturation of at least 3% or an arousal. In version 3 of the manual, the 4% criterion changed from "acceptable" to "optional".

These elements produce the indices that the study report relies on. For automated analysis, one distinction matters: some rules are purely signal-based (a 3% SpO2 drop can be computed deterministically), while others need pattern interpretation (is this epoch N1 already, or still wake?).

The baseline: people disagree too

Before asking how accurate AI is, you have to ask: accurate compared to what? Sleep scoring has no independent ground truth. There is an expert's judgement, and experts differ.

The largest dataset on this comes from the AASM inter-scorer reliability program. Rosenberg and Van Hout (2013) analysed more than 3.2 million decisions by over 2,500 scorers on 1,800 epochs. Mean agreement with the majority score was 82.6%. It was highest for REM, followed by N2 and wake. For N3 it fell to 67.4%, and for N1 to 63.0%.

Respiratory events look similar. In the 2014 analysis, agreement for epochs with any respiratory event was 88.4%, but only 65.4% for hypopneas (κ = 0.57) and 52.4% for central apneas (κ = 0.41).

Bakker et al. (2023) went further and had the same recordings scored by 6, 9 and 12 experts. The share of epochs where every scorer agreed was 46%, 38% and 32% respectively. The authors proposed reporting stage probabilities (a hypnodensity) instead of one label per epoch. That is a useful frame for thinking about AI: a model that says "this epoch is 60% N1, 40% N2" describes the night more honestly than one that pretends to be certain.

What validation studies show

Below are a few papers worth knowing, because they compare an automated system with people on clinical data. This is not a full literature review.

StudyWhat was comparedResultCaveats
Punjabi et al., Sleep 2015 (Somnolyzer)97 recordings scored manually at 4 labs and automaticallyAHI correlation machine vs humans 0.93, between labs 0.92; good agreement on arousal index, total sleep time, sleep efficiencylargest differences in % of N1, N2 and N3; 2007 AASM criteria
Perslev et al., npj Digital Medicine 2021 (U-Sleep)model trained on recordings of 15,660 people from 16 clinical studieson data from a clinic unseen in training, as accurate as the best human expertsleep stages only, no respiratory events
Bakker et al., Sleep 202395 recordings, each scored by 6-12 experts, hypnodensity comparisonICC between automated and manual stage probabilities 0.91also shows how rare full agreement between experts is

Two conclusions recur across these papers. First, at the level of whole-night indices (sleep time, AHI, arousal index), machine and humans agree well. Second, at the level of individual epochs, errors cluster where humans are uncertain: stage transitions, N1, sleep depth, hypopnea type.

The AASM's 2020 position statement on AI in sleep medicine names automated PSG scoring as the most immediate practical application, but sets conditions: a clearly stated population and purpose for the tool, validation on independent data, and transparency towards the lab that uses it.

What is cleared for clinical use

In the US, automated PSG analysis goes through the FDA 510(k) pathway. Two recent examples:

  • SleepStageML (Beacon Biosignals), 510(k) K233438, decided in March 2024. It scores sleep stages from EEG in PSG recordings of adults and is intended to assist a clinician in evaluating sleep. It was the first sleep medicine device cleared with a predetermined change control plan (PCCP), which allows the model to be updated without a new submission, as long as each version passes the agreed tests.
  • SOMNUM v3.0 (HoneyNaps), 510(k) K253390, announced in July 2026. It analyses level 1 PSG recordings of adults aged 22 and over: sleep stages, arousals, limb movements, apneas and hypopneas, classifying apneas as obstructive, central or mixed. Its intended use is under physician supervision, and the physician can edit and delete events.

The AASM also runs its own certification for autoscoring software. The 2023 pilot covered adult sleep staging only, and the first certified product was Sleepware G3 with Somnolyzer. In April 2026 the AASM announced a full-PSG program covering stages, respiratory events, arousals and limb movements, assessed on a multi-site set of recordings. It requires FDA clearance or a pending submission.

In the European Union, software that provides information for diagnostic decisions is usually a medical device. Rule 11 of Annex VIII to the MDR puts it in at least class IIa, which means a notified body is involved. An AI system in such a device is a high-risk system under the AI Act. Regulation (EU) 2026/1744, the AI Omnibus, moved the application date for those obligations for AI in regulated products such as medical devices to 2 August 2028. A tool used only in a research project, with no effect on decisions about an individual patient, is in a different position, but that line has to be drawn deliberately.

This is not legal advice. Classifying a specific piece of software is best done with a medical device regulatory specialist.

What AI does well, and where a human is needed

TaskAutomationHuman role
Desaturations (SpO2 drop ≥ 3% or 4%)signal rule, repeatablecheck sensor artefacts, a slipped oximeter
W, N2, REMusually at the level of inter-scorer agreementreview the hypnogram, especially at transitions
N1 and the N2/N3 boundarymost disagreement, for humans toodecide uncertain epochs
Arousalsdetected from EEG and EMG, sensitive to signal qualityverify, because they affect hypopnea scoring
Apnea and hypopnea typehard, humans agree poorlyclinical judgement in the patient's context
Missing or noisy channelsmany models lose qualitydecide whether the recording can be scored
Populations outside validation (children, neurological disease)no guaranteesout of the tool's scope, manual scoring
Study report and diagnosisoutside the automation's scopephysician only

In practice, a sensible split looks like this: the software does a first pass over the whole night and flags low-confidence epochs and events, and a technologist or physician reviews those first. The time saving comes from not scoring eight hours of signal from scratch, not from nobody looking at the signal anymore.

How we do it: the pipeline for the MT5 Foundation

The MT5 Foundation runs sleep labs in hospitals. Today they record about a hundred nights a month. We delivered the technical side: hardware, deployment, the data route and an NVIDIA Blackwell GPU cluster that sits at the foundation. The full story is in our MT5 Foundation case study.

The pipeline works like this:

  1. Recording. A NOX A1s recorder writes the night to a single EDF+ file. The sample night in our story has 90 channels in the file, 8 hours of recording and about 495 million samples.
  2. Preparation. Anonymisation, filtering (0.3-35 Hz for EEG), a common sampling grid and channel alignment down to the sample.
  3. Describing the night. The pipeline automatically detects airflow cessations, desaturations of at least 3% and arousals, and derives an illustrative hypnogram from the EEG, EOG and activity. We describe events neutrally, without classifying their type.
  4. Windows for the model. The night is cut into 30-second windows on which we train our own model. In each window we mask one channel, such as saturation, and the model has to rebuild it from the rest. To do that, it has to learn how breathing, heart, brain and movement depend on one another.

All of this analysis is illustrative and for research. It is not a physician's report and is not used to diagnose patients. The research questions we are working on are open. We have no results yet and will not announce any before rigorous verification. Patient data never leaves the foundation, because the pipeline, the model and the data sit on hardware in its own infrastructure. Our approach to projects like this is described on the R&D page.

Checklist before rolling out automated scoring

  1. Purpose. Is the tool meant to support clinical reporting, or only research? That determines which clearances you need.
  2. Population. Does the validation cover your patients: age, comorbidities, study type (in-lab PSG or home testing)?
  3. Channels and hardware. Was the model validated on your montage and recorder? What does it do when a channel is missing or noisy?
  4. Rule version. Which version of the AASM manual and which hypopnea criterion (3% or arousal, or 4%)? It has to match what you report.
  5. Your own validation. Compare the software with at least two human scorers on a few dozen of your recordings. Measure machine-human agreement and human-human agreement.
  6. Uncertainty. Does the system show which epochs and events are uncertain? Review should start there.
  7. Audit trail. Can you see what the software did and what a human changed? This matters for quality control and model updates.
  8. Data. Where are recordings processed? If off site, on what legal basis and with what safeguards? If on site, who maintains the server?

Sources