Buyer's guide

Emotion recognition software

A category explainer written by a vendor in it, which means you should read the limitations section with particular attention — it is the part we had least incentive to write.

11 min read
In short

Emotion recognition software infers emotional states from facial expression, voice or physiological signals, typically outputting probabilities across a small set of categories. Commercial options range from browser SDKs at a few euros a month to research platforms at $25,000 per seat. Since February 2025 it has been prohibited in EU workplace and education contexts.

How it works

Nearly every commercial system follows the same five-stage pipeline.

Stage What happens
1. Capture Webcam frames, typically 1–30 fps
2. Face detection Locate faces, extract landmarks (commonly 68 points)
3. Action Unit coding Map landmark configurations to FACS Action Units
4. Classification A model maps Action Units to emotion categories
5. Output A probability distribution, usually summing to 1

Stage 3 is descriptive — it records that a brow lowered. Stage 4 is inferential — it concludes the person is angry. Almost all of the scientific controversy and all of the regulatory weight attaches to stage 4, and vendor accuracy claims usually blur the two.

Most systems classify against the Ekman seven: happiness, sadness, anger, fear, surprise, disgust, contempt. Some use dimensional models instead, placing state on valence and arousal axes rather than in discrete categories.

The main options

Vendor Shape Best for
iMotions Desktop research platform; integrates eye tracking, EEG, GSR Academic and commercial lab research
Affectiva (Smart Eye) Affdex SDK, most-deployed facial coding engine Advertising research, automotive driver monitoring
MorphCast Browser-native HTML5 SDK, client-side Web applications, prototyping, general use
Noldus FaceReader Desktop observational research software Academic behavioural research
Realeyes / Entropik Managed media and advertising testing Ad effectiveness at scale
Emotuit Content engagement analytics for learning Which part of training content loses people

Worth knowing: iMotions licenses the Affdex engine. When you buy facial coding through iMotions, you are frequently buying Affectiva's classifier with a research platform wrapped around it. The market is less differentiated at the classifier layer than the vendor count suggests.

What it costs

Free–€29
Per month, browser SDK tiers (MorphCast)
Published pricing
~$5k
Per year, commercial SDK licensing (Affectiva)
Publicly discussed, scales with usage
$7.5k–25k
Per year, research platform (iMotions academic to commercial per seat)
Publicly discussed, not published

The spread is roughly three orders of magnitude, and it does not track classifier quality. What you are paying for at the top end is hardware integration, study design tooling, participant management, support and academic credibility. The underlying emotion classification is not three thousand times better.

If your requirement is "detect facial expressions in a browser", the cheap end is genuinely adequate. If your requirement is "run a controlled study synchronising eye tracking with EEG", it is not.

What it can measure

Four things it does genuinely well.

Relative change within one person. Did this individual's affect shift between segment A and segment B? With baseline calibration this is reasonably reliable, and it is what most engagement use cases actually need.

Aggregate patterns at scale. Individual error partially cancels across many viewers. "This segment produced more negative affect across 400 people" is far more defensible than any single-person claim.

Timing. Even where the emotion label is uncertain, when an affective shift occurred is often accurate. For content diagnostics, timing is frequently the whole answer.

Presence and gross state change. Whether someone is there, whether they turned away, whether something changed markedly.

What it cannot

Five limitations, covered fully in the limits of facial emotion recognition.

Assert an individual's emotional state at a moment. The inference is probabilistic and context-dependent. Vendor marketing routinely implies certainty the evidence does not support.

Work reliably on spontaneous expressions. Benchmark accuracy figures generally come from posed datasets — actors instructed to display emotions. Real expressions are subtler, briefer, often blended, and frequently absent.

Handle passive viewing well. Facial expression is substantially communicative, something we do at other people. Someone alone at a desk has limited reason to produce one. This is precisely the e-learning scenario.

Perform evenly across populations. Training data composition, cultural display rules, neurodivergence and age all introduce variance. Baseline calibration mitigates this; it does not eliminate it.

Escape the contested premise. The basic-emotion model has been seriously challenged, most prominently by a review led by Lisa Feldman Barrett concluding the mapping between facial configuration and emotional state is far weaker and more context-dependent than the model implies.

The number that reframes the category

Published work puts engagement classification at roughly 91.5% from facial expression alone, rising to 94.6% when behavioural signals are added. Read the other way: the facial layer contributes about three percentage points on top of data collectable from ordinary browser events — and it carries essentially all of the regulatory and cultural cost.

Where you may not deploy it

Since 2 February 2025, EU AI Act Article 5(1)(f) has prohibited AI systems that infer emotions from biometric data in workplace and education institutions. Penalties reach €35 million or 7% of global annual turnover.

Context Status in the EU
Corporate training and L&D ❌ Prohibited
Schools, universities, training providers ❌ Prohibited
Recruitment and interviews ❌ Prohibited
Virtual workspaces and remote work ❌ Prohibited
Market research and ad testing ✅ Lawful, subject to GDPR
Consumer UX research ✅ Lawful, subject to GDPR
Automotive driver monitoring ✅ Safety exception plausibly engages
CE-marked therapeutic use ✅ Medical exception

Three things people get wrong about this:

Consent does not help. Article 5 sets out prohibited practices, not practices permitted with safeguards. There is no consent gateway.

On-device processing does not help. The prohibition concerns the inference, not where it is computed. This is heavily marketed as a privacy answer and it is irrelevant to Article 5.

Both vendor and deployer are in scope. The Act catches placing on the market, putting into service, and use.

Choosing

Work through it in this order.

1. Are you in a prohibited context? EU workplace or education? Then this category is not available to you, and the question becomes what to use instead — usually behavioural signals.

2. What are you actually trying to learn? If the question is "which part of this content loses people", you want content analytics, and facial data is optional. If it is "how did viewers respond emotionally to this advertisement", facial data is the point.

3. Lab or production? Controlled studies with 40 participants need a research platform. Continuous measurement across thousands of learners on their own hardware needs an SDK, and the economics are completely different.

4. What does a false positive cost? For adapting content invisibly, very little — 70% accuracy is fine. For anything touching an individual's record, the error rate is not low enough and the failure modes are not evenly distributed.

5. Will your people accept it? Even where lawful, camera-based monitoring generates resistance. Works councils in much of northern Europe will refuse. Student campaigns against proctoring have been sustained and effective. Factor the cultural cost, not just the licence.


The summary a vendor should not really give you: for the specific job of finding which twelve minutes of a forty-minute module people stop attending to, emotion recognition software is a legitimate technology that is largely unnecessary, frequently unlawful in the relevant contexts, and adds about three percentage points to what browser events already tell you.

Frequently asked questions

How much does emotion recognition software cost?
The range is extreme. Browser SDKs like MorphCast run from a free tier to around €29/month. Affectiva's Affdex SDK has been discussed publicly from around $5,000/year. Research platforms like iMotions run roughly $7,500/year academic to $25,000 per seat commercial. The difference reflects lab hardware integration and support, not classifier quality.
Is emotion recognition software accurate?
On posed, frontal, well-lit expressions it performs well. On spontaneous expressions in real conditions it performs considerably worse, and published engagement classification accuracy of around 91.5% from facial data alone rises only to 94.6% when behavioural signals are added. The underlying basic-emotion model is also scientifically contested.
Can I use emotion recognition software in my company?
It depends where your people are. In the EU, inferring emotions from biometric data in workplace or education contexts has been prohibited since February 2025 under AI Act Article 5(1)(f), with no consent exception. Outside those contexts, and outside the EU, it is generally lawful subject to data protection law.
What is the best emotion recognition software?
There is no single answer because the use cases diverge sharply. For multimodal lab research, iMotions. For a general-purpose browser SDK, MorphCast. For advertising and automotive at scale, Affectiva. For working out which part of your training content loses people, none of them — that is a content analytics problem, and behavioural signals answer it better.

See where your content loses people

Book a walkthrough and we will show you the engagement data on your own content.

Request a demo
Get Started

See what completion rates can't tell you

Find out exactly where your content works, where it fails, and what disengagement looks like before people leave.

Request a Demo