Emotion recognition software infers emotional states from facial expression, voice or physiological signals, typically outputting probabilities across a small set of categories. Commercial options range from browser SDKs at a few euros a month to research platforms at $25,000 per seat. Since February 2025 it has been prohibited in EU workplace and education contexts.
How it works
Nearly every commercial system follows the same five-stage pipeline.
| Stage | What happens |
|---|---|
| 1. Capture | Webcam frames, typically 1–30 fps |
| 2. Face detection | Locate faces, extract landmarks (commonly 68 points) |
| 3. Action Unit coding | Map landmark configurations to FACS Action Units |
| 4. Classification | A model maps Action Units to emotion categories |
| 5. Output | A probability distribution, usually summing to 1 |
Stage 3 is descriptive — it records that a brow lowered. Stage 4 is inferential — it concludes the person is angry. Almost all of the scientific controversy and all of the regulatory weight attaches to stage 4, and vendor accuracy claims usually blur the two.
Most systems classify against the Ekman seven: happiness, sadness, anger, fear, surprise, disgust, contempt. Some use dimensional models instead, placing state on valence and arousal axes rather than in discrete categories.
The main options
| Vendor | Shape | Best for |
|---|---|---|
| iMotions | Desktop research platform; integrates eye tracking, EEG, GSR | Academic and commercial lab research |
| Affectiva (Smart Eye) | Affdex SDK, most-deployed facial coding engine | Advertising research, automotive driver monitoring |
| MorphCast | Browser-native HTML5 SDK, client-side | Web applications, prototyping, general use |
| Noldus FaceReader | Desktop observational research software | Academic behavioural research |
| Realeyes / Entropik | Managed media and advertising testing | Ad effectiveness at scale |
| Emotuit | Content engagement analytics for learning | Which part of training content loses people |
Worth knowing: iMotions licenses the Affdex engine. When you buy facial coding through iMotions, you are frequently buying Affectiva's classifier with a research platform wrapped around it. The market is less differentiated at the classifier layer than the vendor count suggests.
What it costs
The spread is roughly three orders of magnitude, and it does not track classifier quality. What you are paying for at the top end is hardware integration, study design tooling, participant management, support and academic credibility. The underlying emotion classification is not three thousand times better.
If your requirement is "detect facial expressions in a browser", the cheap end is genuinely adequate. If your requirement is "run a controlled study synchronising eye tracking with EEG", it is not.
What it can measure
Four things it does genuinely well.
Relative change within one person. Did this individual's affect shift between segment A and segment B? With baseline calibration this is reasonably reliable, and it is what most engagement use cases actually need.
Aggregate patterns at scale. Individual error partially cancels across many viewers. "This segment produced more negative affect across 400 people" is far more defensible than any single-person claim.
Timing. Even where the emotion label is uncertain, when an affective shift occurred is often accurate. For content diagnostics, timing is frequently the whole answer.
Presence and gross state change. Whether someone is there, whether they turned away, whether something changed markedly.
What it cannot
Five limitations, covered fully in the limits of facial emotion recognition.
Assert an individual's emotional state at a moment. The inference is probabilistic and context-dependent. Vendor marketing routinely implies certainty the evidence does not support.
Work reliably on spontaneous expressions. Benchmark accuracy figures generally come from posed datasets — actors instructed to display emotions. Real expressions are subtler, briefer, often blended, and frequently absent.
Handle passive viewing well. Facial expression is substantially communicative, something we do at other people. Someone alone at a desk has limited reason to produce one. This is precisely the e-learning scenario.
Perform evenly across populations. Training data composition, cultural display rules, neurodivergence and age all introduce variance. Baseline calibration mitigates this; it does not eliminate it.
Escape the contested premise. The basic-emotion model has been seriously challenged, most prominently by a review led by Lisa Feldman Barrett concluding the mapping between facial configuration and emotional state is far weaker and more context-dependent than the model implies.
Published work puts engagement classification at roughly 91.5% from facial expression alone, rising to 94.6% when behavioural signals are added. Read the other way: the facial layer contributes about three percentage points on top of data collectable from ordinary browser events — and it carries essentially all of the regulatory and cultural cost.
Where you may not deploy it
Since 2 February 2025, EU AI Act Article 5(1)(f) has prohibited AI systems that infer emotions from biometric data in workplace and education institutions. Penalties reach €35 million or 7% of global annual turnover.
| Context | Status in the EU |
|---|---|
| Corporate training and L&D | ❌ Prohibited |
| Schools, universities, training providers | ❌ Prohibited |
| Recruitment and interviews | ❌ Prohibited |
| Virtual workspaces and remote work | ❌ Prohibited |
| Market research and ad testing | ✅ Lawful, subject to GDPR |
| Consumer UX research | ✅ Lawful, subject to GDPR |
| Automotive driver monitoring | ✅ Safety exception plausibly engages |
| CE-marked therapeutic use | ✅ Medical exception |
Three things people get wrong about this:
Consent does not help. Article 5 sets out prohibited practices, not practices permitted with safeguards. There is no consent gateway.
On-device processing does not help. The prohibition concerns the inference, not where it is computed. This is heavily marketed as a privacy answer and it is irrelevant to Article 5.
Both vendor and deployer are in scope. The Act catches placing on the market, putting into service, and use.
Choosing
Work through it in this order.
1. Are you in a prohibited context? EU workplace or education? Then this category is not available to you, and the question becomes what to use instead — usually behavioural signals.
2. What are you actually trying to learn? If the question is "which part of this content loses people", you want content analytics, and facial data is optional. If it is "how did viewers respond emotionally to this advertisement", facial data is the point.
3. Lab or production? Controlled studies with 40 participants need a research platform. Continuous measurement across thousands of learners on their own hardware needs an SDK, and the economics are completely different.
4. What does a false positive cost? For adapting content invisibly, very little — 70% accuracy is fine. For anything touching an individual's record, the error rate is not low enough and the failure modes are not evenly distributed.
5. Will your people accept it? Even where lawful, camera-based monitoring generates resistance. Works councils in much of northern Europe will refuse. Student campaigns against proctoring have been sustained and effective. Factor the cultural cost, not just the licence.
The summary a vendor should not really give you: for the specific job of finding which twelve minutes of a forty-minute module people stop attending to, emotion recognition software is a legitimate technology that is largely unnecessary, frequently unlawful in the relevant contexts, and adds about three percentage points to what browser events already tell you.
Frequently asked questions
How much does emotion recognition software cost?
Is emotion recognition software accurate?
Can I use emotion recognition software in my company?
What is the best emotion recognition software?
See where your content loses people
Book a walkthrough and we will show you the engagement data on your own content.