The product review ends with no obvious conflict. One director says the new onboarding flow "looks much cleaner." Another asks when it could ship. Two people say nothing before the call runs out of time.
Later, an AI analysis labels the room's reaction as positive.
That sounds plausible. It is also more certainty than the meeting produced. "Cleaner" may be praise for the design, not support for the launch plan. A question about timing may signal enthusiasm, concern, or simple curiosity. Silence may mean agreement, distraction, caution, or that nobody created space to speak.
This is the trap in meeting sentiment analysis. A useful pattern detector can quietly become a narrator of other people's inner states. Once its label enters a recap, "mostly positive" can harden into "the group agreed," then into a roadmap commitment no one actually made.
The better goal is not to make AI guess feelings more confidently. It is to make the evidence of reaction easier for people to inspect.
Sentiment is a clue, not a verdict
The idea is timely. On September 24, Google published a list of five AI agents it recommends executives build first. One is a "Sentiment Analyst" that reviews meeting transcripts with natural-language processing to gauge attendee tone and reaction. Google says it can highlight moments of engagement and confusion so follow-up is more targeted. Google Workspace
There is real value here. Across dozens of interviews, customer calls, or town halls, software can help a person notice recurring objections, unusually active topics, repeated questions, and changes in vocabulary. Those patterns can tell a leader where to look.
Trouble begins when a directional signal is treated as a psychological fact. Conversation is full of language that performs several jobs at once. "That's ambitious" could be encouragement, skepticism, or a polite warning. "I can work with that" could mean genuine support or reluctant compliance. A transcript may preserve the words while losing the pause, facial expression, prior disagreement, power difference, and private calculation around them.
Even research systems designed specifically for conversational emotion face this problem. A 2025 study on uncertainty in emotion recognition describes biased predictions and poor calibration, including a tendency for classifiers to favor categories such as neutral. Cambridge University Press Another study notes that the same sentence can carry opposite emotional tones in different conversational contexts, which is one reason researchers combine text with audio and surrounding dialogue. Scientific Reports
This does not make analysis useless. It tells us what the output should be: a prompt for inquiry, not a verdict about a person.
A transcript records language, not consent
Before analyzing a meeting, separate four things that summaries often collapse:
- Expression: the words someone used, including questions, objections, praise, and conditions.
- Behavior: an observable act, such as volunteering, deferring, changing a proposal, or declining ownership.
- Interpretation: a plausible explanation for the expression or behavior.
- Commitment: an explicit decision, approval, task, or promise.
Only the last category should move work by itself. The others may inform the next question.
Consider the director who asks, "When could this ship?" The expression is a timing question. The behavior is active engagement with the proposal. One interpretation is excitement; another is concern about feasibility. There is no commitment until the director says something like, "I approve the launch for October 12," or accepts a clearly stated decision rule.
Good meeting records preserve this distance. Weak ones skip from expression to commitment because the middle feels obvious.
Build a reaction ledger
A reaction ledger is a compact record of what happened around an important proposal. It gives AI a narrower job than assigning a mood to the room.
For each consequential topic, capture four fields:
- Proposal: What exactly was introduced?
- Observed response: What did participants say or do that is visible in the record?
- Open interpretation: What might the response mean, and what remains ambiguous?
- Confirmation needed: Which direct question, owner, or deadline will resolve the ambiguity?

For the onboarding review, the ledger might read:
| Proposal | Observed response | Open interpretation | Confirmation needed |
|---|---|---|---|
| Ship the redesigned flow on October 12 | Design called "cleaner"; one timing question; no explicit objections; two participants did not comment | Visual direction may be supported; launch readiness was not confirmed | Ask each function to confirm readiness or name a blocker by Thursday |
Notice what the ledger refuses to claim. It does not call the silent participants disengaged. It does not translate a compliment into approval. It also does not ignore the positive signal. It preserves that signal at the level the evidence supports.
This structure works for more than executive meetings. In a customer interview, it keeps "that could be useful" separate from purchase intent. In a hiring debrief, it prevents warmth or conversational ease from substituting for job evidence. In a retrospective, it distinguishes frustration with one incident from opposition to the broader process.
Ask before the room disperses
The cheapest ambiguity to resolve is the one you catch while everyone is still present.
Near the end of a consequential topic, the facilitator can ask three short questions:
- "What part of this proposal do you support?"
- "What concern or condition have we not recorded?"
- "What are you personally committing to, if anything?"
These questions give participants distinct ways to respond. A person can like the direction while naming a dependency. Someone can understand the plan without owning a task. A quiet specialist can add a condition without having to challenge a supposed consensus.
Do not demand emotional disclosure. "How do you feel about this?" may be appropriate in some teams, but it can also pressure people to make private reactions public. Questions about support, concern, evidence, and commitment are usually easier to answer and more useful for the work.
If the meeting is too large for a spoken round, use a brief written check. Ask participants to choose support, support with condition, need more evidence, or object, then require one sentence of context for anything beyond simple support. The labels describe the proposal's status, not anyone's personality.
Use claims that can be corrected
An AI-generated meeting analysis should make it easy for a participant to say, "That is not what I meant."
Prefer language tied to evidence:
Three participants raised implementation questions. One person supported the direction with a staffing condition. No explicit approval was recorded.
Avoid language that assigns hidden states:
The team was excited, although two members seemed resistant.
The first statement can be checked against the record. The second asks readers to trust an interpretation of emotion and turns ambiguity into a reputation. That matters when recaps travel to managers, performance discussions, account records, or future meetings where the people described are not present.
Also keep aggregate patterns away from individual judgments. It may be reasonable to observe that questions about migration appeared in eight customer calls. It is far more consequential to label a particular customer "negative" or an employee "unenthusiastic" based on a small slice of conversation.
Let AI point to the next question
The most responsible use of sentiment analysis is diagnostic. It can surface a cluster of cautious phrases, an abrupt change in participation, repeated requests for evidence, or a mismatch between lively discussion and the absence of commitments. Each is a reason to investigate.
Ask the system to show its work: which utterances drove the pattern, whose voices are missing, what alternative readings are plausible, and which direct question would reduce uncertainty. If it cannot point back to observable evidence, its label should not enter the decision record.
A meeting copilot is especially useful when it helps in the moment, before an uncertain reaction gets rewritten as consensus. Caspi provides live recap, suggested questions, contextual chat, proactive flags from connected tools, post-call action items, and persistent meeting memory. Those capabilities can support a reaction ledger by preserving what was actually said, prompting a clarifying question, and carrying explicit commitments forward without pretending to know what silence meant.
AI can help us notice where a conversation deserves attention. The human responsibility is to ask, confirm, and record the answer at the level of certainty the room actually earned.