The million-dollar slide: what your sales transcripts are missing
How multimodal video understanding changes what companies can extract from their video archives.

Picture this: your top sales rep just closed a major deal after showing a competitor comparison slide that highlighted your 40% cost advantage. The prospect's eyes lit up, they leaned forward, and said "that changes everything."
Your current call analysis system recorded the conversation. It captured the words "competitor," "pricing," and "interesting." But it completely missed the moment that actually closed the deal — the visual proof on the slide that made your value proposition undeniable.
This is the million-dollar problem with speech-only video analysis. While your transcripts capture what was said, they're blind to what was shown. And in sales, what you show often matters more than what you say.
The visual intelligence gap in enterprise video
At Cloudglue, we've been tackling this exact challenge. Our research, accepted to the KDD 2025 Agentic AI for Enterprise workshop, reveals just how much traditional speech-only analysis leaves on the table.
| Query type | Speech-only | Multimodal | Gain (percentage points) |
|---|---|---|---|
| Speech queries | 80% | 90% | +10 |
| Visual queries | 20% | 100% | +80 |
| Cross-modal | 60% | 100% | +40 |
| Overall | 53.3% | 96.7% | +43.4 |
Evaluation scope: 15 queries (five per modality) over a two-hour collection of sales and product-demo videos. Human annotators scored answers as correct, partially correct, or incorrect. These are the study's correctness scores; the gain column shows the difference in percentage points. Read the evaluation and limitations in the VideoMCP paper (PDF).

For visual queries specifically, speech-only gets just 20% right, while multimodal achieves 100%. Think about what that means for your business. When your sales manager asks "which slides generated the strongest reactions this quarter?" or "what pricing objections came up on screen versus in conversation?" — speech-only tools literally cannot answer these questions.
What you're missing right now
Every day, your organization generates hours of valuable video content:
- Sales calls with competitor mentions on shared screens
- Product demos showing specific UI interactions that drive decisions
- Training sessions where the presenter's slides contain the actual process steps
- Customer success reviews with performance dashboards that tell the real story
Traditional transcription captures the presenter saying "here's our Q3 performance…" but misses the actual metrics displayed. It logs "let me show you this feature" but ignores the workflow demonstration that convinced the customer. Your most valuable insights are hiding in plain sight — literally.
The science behind multimodal video intelligence
Our VideoMCP system, now available as the Cloudglue MCP Server, doesn't just transcribe — it sees your videos. It combines:
- Automatic speech recognition for spoken content
- Optical character recognition for text on slides, screens, and documents
- Visual scene analysis for understanding context and reactions
- AI agent orchestration to synthesize insights across all modalities
The result is an AI that can answer questions like:
- "Show me all mentions of [competitor X] along with the slides that were displayed"
- "What metrics were discussed versus what was actually shown on screen?"
- "Which demo sections generated visible customer engagement?"
- "Find compliance topics that were both spoken about and displayed in documentation"
From research to revenue impact
The implications extend far beyond technical metrics. Early adopters report sales teams identifying winning demo patterns by analyzing both presentation content and customer visual reactions; compliance teams verifying that required topics appear in both speech and supporting documentation; training organizations linking spoken instructions to actual UI demonstrations; and customer success teams correlating verbal feedback with on-screen performance data.
While your competitors rely on transcripts that capture half the story, you could be leveraging the complete picture. Every slide that swayed a prospect. Every UI interaction that demonstrated value. Every visual proof point that closed a deal.
