Published research identifies tab-switching as the single strongest predictor of disengagement in online courses, ahead of self-regulation and satisfaction measures. Adding facial expression analysis to behavioural signals improves engagement classification accuracy from roughly 91.5% to 94.6% — a gain of about three percentage points.
The two headline results
The 2024 ScienceDirect study examined predictors of disengagement in online courses and found cyberloafing — switching away from the learning content — outranked both self-regulation and satisfaction measures.
Separately, the OUCI/DNTB work on multimodal classification reports accuracy improving from 91.5% to 94.6% when behavioural signals are combined with facial expression.
Reading the fusion study correctly
The 91.5% → 94.6% figure is quoted constantly in this industry, almost always to demonstrate that multi-modal approaches are superior. That reading is correct and incomplete.
What it also establishes is the size of the facial contribution: roughly 3.1 percentage points, on top of signals available from ordinary browser events.
Turn the comparison round and it reads differently. If behavioural signals alone approach the low 90s — which the cyberloafing result makes plausible — then the architecture most of this market sells is a large amount of regulatory, technical and cultural complexity in exchange for a marginal gain.
These are separate studies with different populations, tasks and disengagement definitions. Treating them as a clean head-to-head overstates what either establishes individually. The defensible claim is directional: behavioural signals carry most of the predictive weight, and facial analysis is an increment on top rather than the foundation. Anyone presenting this as a precise trade-off — including us — is over-reading the evidence.
Why behavioural signals do so well
Three reasons, and they are worth understanding because they suggest the finding will hold rather than being an artefact.
They are unambiguous. A tab-switch is a discrete, recorded event. There is no classification step, no model, no confidence interval. It either happened or it did not. Facial analysis requires detection, landmark extraction and classification, each contributing error.
They are behaviour, not inference. Switching away from content is disengagement, rather than being evidence from which disengagement is inferred. You are measuring the thing, not a correlate of it.
Faces do very little during passive viewing. This is the most under-discussed fact in the field. Facial expression is substantially communicative — something we do at other people. A learner alone at a desk watching a module has limited reason to produce one. They can be thoroughly frustrated and display almost nothing.
That third point matters especially for e-learning, because passive viewing is the dominant mode. The scenario where facial analysis has least signal is precisely the scenario it is most often sold into.
Where facial analysis genuinely adds
Having spent three sections on the limitation, the honest other side.
Distinguishing confusion from boredom. Behavioural data tells you attention dropped. It cannot tell you whether the learner was lost or bored, and those need opposite fixes — one wants better explanation, the other wants cutting. Affective data offers a hypothesis where behavioural data offers none.
Passive video specifically. In long-form video where correct behaviour is stillness, behavioural signals are thin. There is little to measure when the right thing to do is sit still. This is the case where the facial layer contributes most.
Timing precision. Even where the emotion label is uncertain, when an affective shift occurred is often accurate — and for content diagnostics, timing is frequently the whole answer.
Aggregate patterns at scale. Individual classification error partially cancels across many viewers. "This segment produced more negative affect across 400 learners" is a considerably more defensible claim than anything about one person.
What this means commercially
For the buyer, the question is not "which signal is better" but "what am I trading."
| Behavioural only | With facial analysis | |
|---|---|---|
| Predictive accuracy | Most of it | +~3 percentage points |
| Camera required | No | Yes |
| Biometric data | None | Yes |
| GDPR special category | No | Yes |
| EU workplace/education | ✅ Lawful | ❌ Prohibited |
| Works council | Routinely approved | Frequently refused |
| Confusion vs boredom | Not available | Available |
| Passive video sensitivity | Weak | Stronger |
For corporate L&D and education in the EU, the bottom-left cell decides it — the trade is not available, because Article 5(1)(f) prohibits it. For market research and media testing, where facial analysis is lawful and passive viewing dominates, the increment is worth having.
This is why Emotuit ships both configurations rather than picking one. It is also why Signals — behavioural only — is the default recommendation for learning use cases, despite being the configuration that uses less of our own technology.
Frequently asked questions
Does this mean facial analysis is useless?
Why do vendors lead with facial analysis then?
Is 3 percentage points not worth having?
Are these studies directly comparable?
See where your content loses people
Book a walkthrough and we will show you the engagement data on your own content.