Using Voice — Unique AI Documentation

Using Voice

Overview

Voice is offering users a seamless way to convert spoken audio into real-time transcriptions. This capability lets users speak their prompts instead of typing, enhancing accessibility and interaction within AI-powered experiences.

If you’re unable to access certain features or sections of this article, it’s possible that your firm doesn’t have access or hasn’t upgraded to the latest version. Please reach out to your internal support team for further assistance.

Who it’s for

Benefits

Speed & Efficiency
Speaking is often faster than typing, allowing users to input prompts more quickly — especially helpful for long or complex queries.

Hands-Free Operation
Ideal for multitasking or mobile workflows where typing may be impractical or disruptive.

Natural Language Input
Encourages more conversational, free-form prompts, which align well with how GPT-based systems interpret intent.

Accessibility
Enables users with motor impairments, temporary injuries, or vision difficulties to interact more easily with AI systems.

Reduced Cognitive Load
Users can focus on articulating their ideas instead of editing text as they type.

Example Use cases

Financial Analysts on the Move
Analysts commuting or walking between meetings can speak questions like “What’s the impact of today’s Fed announcement on bond yields?” into the prompt box without stopping to type.

Executive Summarization
A manager uses voice to say: “Summarize Q1 performance highlights from our internal reports and flag any anomalies.”

Customer Service Triage
Agents can dictate incoming customer questions directly into the system — useful for hands-busy environments or for real-time escalation.

Step-by-Step Guide

Privacy Note: Voice data is not stored. It is transcribed in real time and discarded immediately after transcription.

If no toggle is present, refer to the Voice Administration guide to set up the landscape correctly.

Documentation for Infrastructure is available at Voice Infrastructure.

Tips & Tricks

Tap, Speak, Send
Use the microphone icon to quickly input long or complex prompts hands-free. Especially useful for mobile workflows or when referencing multiple variables aloud.

Speak Naturally
The system is optimized for natural language. Don’t worry about being overly formal — just talk as if you're asking a colleague.

Pause for Clarity
Brief pauses help improve transcription accuracy and let the system segment your speech correctly (useful for multi-part prompts).

Combine With Keyboard if Necessary
Speak the main part of your prompt, then use the keyboard to fine-tune or add parameters (e.g., “... but exclude any data before 2020”).

If you don’t see the microphone toggle, please double check the following:

  1. You’re using a supported browser.
  2. You’ve granted microphone permissions.
  3. Your device has a working microphone.
  4. The feature has been set up successfully as outlined in Voice Administration.

Limitations

We currently support four languages: English, German, French, and Italian. English and German work better, with fewer mistakes.

French and Italian have some challenges, like:

Our solution uses Microsoft’s Azure Text-to-Speech, so we're limited by what it can do. We'll update and improve as new features become available.

Real-time streaming is technically impossible due to the network latency and processing of the AI services. The average stream response time is between 1~5 seconds (5G, Wifi) depending on the network of the browser/user.