Meeting engagement analytics measures whether participants are attending to a virtual session and where collective attention breaks. The credible approaches are post-session analysis of the recording, behavioural telemetry for browser-delivered content, and designed-in interaction. Inferring emotions from participants' faces is prohibited in EU workplace contexts.
Why meetings are unmeasured
Organisations run all-hands, town halls, webinars and virtual training constantly, at enormous aggregate cost in salaried hours, and measure almost none of it.
The reasons are structural:
No native instrumentation. Zoom removed attention tracking in 2020. Teams and Meet never had an equivalent. The conferencing layer offers attendance and little else.
Surveys arrive too late and measure the wrong thing. A post-session form captures recalled feeling, dominated by how the session ended. It cannot tell you the room went quiet at minute 22.
Everyone knows it is awkward. Measuring whether people paid attention to your CEO sits close enough to surveillance that most organisations decline to try.
The result is that the highest-cost, lowest-measured activity in the business runs on vibes.
What conversation intelligence measures
Gong, Avoma, Chorus and similar tools have built a substantial category here, and it is worth being clear about what they do — because it is frequently conflated with engagement.
They analyse the speaker: transcript, talk-to-listen ratio, filler words, monologue length, question rate, topic coverage, next steps.
That is genuinely valuable for sales coaching, and it is not audience engagement. A rep with a textbook 43% talk ratio, good question rate and clean next steps can still have lost the buyer at minute six. The transcript records what was said. It does not record whether anyone was listening.
What audience engagement measures
The complementary question: was the room with you, and where did you lose it?
This is a different measurement problem with different signals:
| Conversation intelligence | Audience engagement | |
|---|---|---|
| Subject | The speaker | The participants |
| Source | Audio and transcript | Behaviour and attention |
| Answers | What was said, how | Whether anyone was attending |
| Output | Coaching on delivery | Which moment lost the room |
| Vendors | Gong, Avoma, Chorus | Sparse |
The right column is thin, which is precisely why it is an opportunity rather than a crowded market.
The three credible methods
1. Post-session analysis of the recording
If the session was recorded, the raw material for a good analysis already exists.
Analysing afterwards has an important property beyond convenience: it is diagnostic rather than supervisory. Nobody is watching a live attention dashboard. You are reviewing a session to improve the next one, which is a materially easier conversation with participants and with a works council.
You get: per-participant engagement across the timeline, the moments where collective attention dropped, and — critically — the transcript aligned to those moments, so you can see what was being said when the room left.
2. Behavioural telemetry for browser content
Where training is delivered as browser content rather than through a conferencing client — an LMS module, a deck, a recorded video — instrument the content itself.
This gives per-section resolution: tab-switching, window focus, dwell per segment, scroll velocity, interaction latency, replay, abandonment. All from ordinary browser APIs, no camera, no biometric data.
3. Designed-in interaction
Build moments requiring a response: polls at intervals, chat prompts, breakout tasks with outputs, Q&A.
Low-tech, and it has two real advantages. It is unambiguous — a poll response is direct evidence of presence — and it is completely uncontroversial. Nobody objects to being asked a question.
Its weakness is sampling: it tells you about the moments you asked, not the gaps between them.
Metrics worth tracking
| Metric | Signal quality | Notes |
|---|---|---|
| Collective attention curve | ★★★★★ | Where the room dropped, per minute |
| Drop-off position | ★★★★★ | When people left the session entirely |
| Poll response rate | ★★★★ | Direct evidence, but sampled |
| Q&A volume and timing | ★★★★ | Timing matters more than count |
| Chat activity over time | ★★★ | Biased toward extroverts and juniors |
| Focus-adjusted dwell | ★★★★ | For browser content only |
| Attendance duration | ★★ | Baseline only |
| Camera-on rate | ★ | Reflects culture, not attention |
| Reaction usage | ★ | Culturally variable, easily gamed |
The single most useful output is the collective attention curve — engagement per minute across the session, aggregated across participants. Individual data is rarely what you need and always what causes trouble. "Attention dropped 40% between minutes 18 and 24" is a fact about your content. "Sarah was distracted" is a fact about Sarah, and acting on it is a management problem rather than a content one.
The compliance boundary
Since 2 February 2025, EU AI Act Article 5(1)(f) has prohibited AI systems that infer emotions from biometric data in workplace and education contexts. Commission guidance reads "workplace" to include virtual workspaces and remote work.
A video call with your employees is therefore squarely within the prohibited context. Facial emotion analysis of participants is not available to you in the EU, regardless of consent, contract wording or on-device processing.
| Approach | EU workplace |
|---|---|
| Attendance and duration | ✅ Fine |
| Poll, chat and Q&A data | ✅ Fine |
| Behavioural telemetry in browser content | ✅ Fine |
| Post-session behavioural analysis | ✅ Fine |
| Facial emotion inference | ❌ Prohibited |
| Voice emotion analysis | ❌ Prohibited |
GDPR still applies to everything in the green rows. You will need a lawful basis, and a DPIA for systematic monitoring.
Report at session level rather than participant level and most of the difficulty dissolves. "Attention dropped 40% at minute 18" needs no individual identification, no per-person profile and no uncomfortable conversation — and it is the finding you actually wanted.
Where to start
Pick one recurring session that matters. The monthly all-hands, the mandatory induction, the flagship webinar. Something that repeats, so improvements compound.
Analyse one recording. Not a programme, not a rollout. One session, to see whether the data tells you anything you did not already know. It usually does, and it is usually a specific segment.
Change one thing. Restructure the segment where attention dropped. Cut it, move it, or change its format.
Measure the next one. A before-and-after on the same recurring session is the most persuasive evidence available, because the audience and context are held constant.
Then decide whether to scale. If one session produced a change worth making, the case for instrumenting more makes itself. If it did not, you have spent one analysis finding out.
Frequently asked questions
What is the difference between meeting intelligence and engagement analytics?
Can I measure engagement in a Zoom meeting?
Is talk-time a good engagement metric?
Is camera-on rate a useful signal?
Start with a session you already recorded
Upload one recording and see the engagement timeline, drop-off points and transcript alignment.