
ElevenLabs announced on LinkedIn that ElevenAgents now supports multimodal inputs, allowing AI agents to process images, PDFs, audio messages, contact cards, and location pins alongside voice and text. The update addresses a core limitation in conversational AI: customers communicate through multiple formats, but agents have historically been confined to text and voice only.
The company, which recently launched ElevenAgents for Support to power customer service voice AI, is now extending those agents to handle the full range of media customers send during support interactions.
WhatsApp and Web Widget Integration
On WhatsApp, ElevenAgents can now interpret photos and documents sent by customers. According to the announcement, an agent can analyze a photo of a router, read indicator lights, run diagnostics, and schedule a technician visit without human escalation. The platform also processes PDFs, audio messages, contact information, and location pins shared through WhatsApp threads.
The ElevenAgents web widget now accepts image and PDF uploads directly in the chat interface. ElevenLabs cited use cases including proof of address documents, medical records, and insurance claims, all processed in real time within the same conversation.
Unlike static file upload forms, these agents maintain conversational context across multiple attachments and modalities. A customer can describe an issue, upload a supporting document, then continue the conversation based on what the agent extracted from the file.
Cross-Channel Context Retention
ElevenLabs emphasized that agents maintain full context across channels and input types. An agent on a voice call can send a text message mid-conversation to confirm an appointment or share a document, then process the customer's response, whether it arrives as text, voice, or an image.
This cross-channel continuity differentiates the platform from traditional chatbots or IVR systems, which typically segment interactions by channel. A customer who starts a voice call and later sends a photo via WhatsApp won't need to re-explain their issue; the agent already has the conversation history and can integrate the new input into its next action.
The update builds on ElevenLabs' positioning as a full-stack conversational AI platform. The company became the first AI voice platform to secure agent insurance and has been iterating rapidly on agent capabilities, including A/B testing through its Experiments feature.
Implications for Customer Support Workflows
Multimodal support removes a common bottleneck in automated customer service. When agents can't process attachments, customers must either describe visual information verbally (often inaccurately) or escalate to a human agent. By reading files and images directly, ElevenAgents can resolve issues that previously required handoffs.
For businesses, this means fewer escalations and faster resolution times. A customer uploading a utility bill for account verification, sending a photo of a damaged product, or sharing a map location for delivery instructions can now complete these actions without leaving the agent conversation.
ElevenLabs has not disclosed pricing changes associated with the multimodal features or provided a timeline for availability across all ElevenAgents tiers. The company's recent $500M Series D funding and expansion into government and enterprise sectors suggest continued investment in agent capabilities.
The move aligns ElevenLabs more closely with full-service customer engagement platforms like Intercom and Crisp, which already support file sharing in live chat. However, ElevenLabs is layering multimodal AI interpretation on top of that infrastructure, enabling agents to act on the content of those files rather than simply routing them to human agents.
Sources
1 checkedHow we cover tool news: Create With's tool desk drafts these reports with AI from the sources listed above and checks them against those sources before publishing.




