AI transcription for the next step.
Speak a request, bring in a recording, or send a voice note through a supported connection. Ghost turns the speech into text you can use in your work—and can turn a written reply into a speech file when you ask.
Speech to text · Text to speech · Private beta for Mac
A spoken thought, with a next step.
- Input
- “Keep the Atlas rollout question in a note. I need to check the setup wording.”
- Review
- Read the transcript and clarify the task before acting on it.
- Continue
- Save the note, create a follow-up, or ask for a spoken summary.
From a spoken thought to a reviewed next step.
Speak a request or bring in a permitted recording. Review the transcription, choose what should happen next, and request a speech file when listening is useful.
Start the recording when you want to speak.
ghost’s Mac composer requests microphone access and shows the incoming transcript. You initiate the session; this is not continuous background listening.
“Keep the Atlas setup question in a note. I need to review the wording before the next meeting.”
- Control
- Start and stop the recording yourself
- Review
- Inspect the words before using them as the request
Read the recording as text.
An audio or video file follows the approved file workflow into the sandbox. The configured service returns text, with timestamps or speaker labels when supplied.
Choose what the transcript should become.
Ask ghost for a useful note or a specific follow-up. Keep the original statement separate from the summary and do not infer a deadline or owner that the recording never established.
- Note
- Atlas setup — wording to review
- Follow-up
- Check the setup explanation before the next review
- Open detail
- Confirm the deadline rather than inventing it
Request a speech file from the text.
ghost can generate a WAV briefing using a supported built-in voice. Generation runs in the agent backend; choosing where to deliver the file is a separate action.
- Text
- The reviewed short briefing
- Voice
- A supported language and built-in voice
- Output
- A WAV file prepared for explicit delivery
Illustrative transcript and output only. The walkthrough does not record your microphone, upload audio, transcribe media, create speech, or send a message.
What is AI transcription?
AI transcription converts speech from an audio stream or recording into written text. Ghost uses it for user-started voice input, supported voice messages, and audio or video files prepared for processing. The transcript becomes input for the same assistant that works with notes, tasks, meetings, and research. Speech generation is the reverse direction: creating audio from text you provide.
Speak the request. Keep control of the message.
Start voice recording in Ghost’s Mac composer when speaking is easier than typing. The app requests microphone access, shows transcription state, and displays the incoming transcript as the session progresses. Recording starts through your action; this is not an always-listening assistant.
Stop the recording, read the captured words, and make the request clear. Names, numbers, specialized terms, background noise, and overlapping speech can affect the result. A visible transcript is useful because you can inspect what Ghost will work from.
Voice input uses the configured transcription service. Its availability depends on the required connection and permissions; the existence of a microphone button is not a promise of offline transcription.
- Start
- You initiate the recording and grant microphone access.
- See
- The composer shows listening or processing state and the transcript.
- Send
- Use the captured text as your message to Ghost.
Transcribe audio and video files.
Bring a recording into the supported file workflow, then ask Ghost to transcribe it. The transcription tool reads a prepared audio or video file from the sandbox workspace and sends it to the configured speech-recognition service.
The result is a text transcript. When the service returns timestamps or speaker labels, those can be retained in the output; a recording without reliable speaker information should not be presented as a verified account of who said what.
For a file on your Mac, the upload step requires approval before processing. Use recordings you have permission to share, and review important statements against the source audio. This page covers file and message transcription; live meeting notes have their own capture and insight workflow.
Preserve the speech before summarizing it.
- Source
- A permitted audio or video recording prepared for transcription.
- Transcript
- Speech text, with supplied timestamps or speaker labels where available.
- Check
- Confirm names, numbers, and consequential statements against the recording.
Turn voice notes into something you can use.
On a linked Telegram conversation, Ghost can receive a voice message and pass the transcription into the conversation. Voice notes, audio messages, and video notes are supported when the required transcription service is enabled.
Tell Ghost what should happen after transcription. Save the useful substance in an editable note, extract a follow-up for To-dos, or use the text to frame a research question. These are explicit next steps, not an assumption that every recording needs a task.
Keep the transcript and interpretation distinct. A summary condenses the speech; a task describes what should happen next. Neither should quietly add a deadline, owner, or promise that the original recording did not establish.
- Capture
- Read the words returned from the recording or supported voice message.
- Interpret
- Ask for a summary, structured note, or question to investigate.
- Act
- Create the specific follow-up you requested, with uncertain details left open.
Text to speech, when listening is useful.
Ask Ghost to turn a written passage into a speech file: a concise briefing, a draft explanation, or a summary you want to listen to. The speech tool generates a WAV file using one of its built-in voices, then hands the file to the delivery workflow.
The current implementation lists English, French, German, Italian, Portuguese, and Spanish, with language-specific built-in voice choices. Availability depends on the speech runtime being configured. This does not add voice cloning, impersonation, or an outbound calling service.
Speech is generated in Ghost’s agent backend and saved in the sandbox for delivery. That is not the same as saying all audio remains on your Mac. Creating the file and sending it to a person or channel are separate actions.
- Input
- The text you want spoken, plus a supported language and voice.
- Output
- A generated WAV speech file—not a live phone conversation.
- Delivery
- Review the result and choose the supported destination explicitly.
Product details
Voice, at a glance.
Supported work, useful outputs, and what to check before you start.
- Input
- Start microphone input in Ghost, or transcribe a supported audio or video file prepared in the sandbox. File transcription needs the configured speech service.
- Output
- Use the transcript as text for notes and tasks. Speech generation is a separate operation that creates a WAV file using a supported built-in voice.
- Delivery
- Creating speech does not send it to anyone. Audio delivery is a separate step, and transcription quality or speaker labels depend on the recording and service output.
One assistant. Connected work.
Ghost is a personal AI assistant for Mac. Voice is one way to bring information into Ghost and receive an answer from the same assistant. Transcripts can feed Notes, To-dos, Research, and Meetings-related work; those connections are deliberate parts of your task, not separate voice-only products.
Speak an observation after a review. Check the transcript, save the useful details as a note, create the follow-up task, and ask for a short spoken briefing before returning to the work.
Questions about voice and transcription.
Is Ghost always listening?
No. Voice input in the Mac composer starts when you initiate recording and grant microphone access. This page does not describe an always-on wake-word assistant.
Can Ghost transcribe an audio or video file?
Yes, through the supported file workflow. The file must be available in the sandbox workspace, and the transcription service must be configured. Uploading a local file requires approval. Review important transcript details against the original.
Can it identify speakers and include timestamps?
File transcription can preserve speaker labels and timestamps returned by the service. Their availability and accuracy depend on the source and transcription path. An unknown speaker should remain unknown rather than being assigned an invented identity.
Does it work with Telegram voice notes?
The Telegram bridge supports incoming voice notes, audio, and video notes when the required transcription service is enabled. This is configuration-dependent and is not a claim that every messaging app supports the same workflow.
Can Ghost read text aloud or clone a voice?
Ghost can generate a WAV speech file from text using its supported built-in voices. The speech tool does not accept an arbitrary reference voice or provide voice cloning. It is not an outbound phone-call agent.
Does all voice processing happen locally on my Mac?
No such guarantee is made. Mac voice input uses a configured speech-recognition connection, file transcription sends prepared media to the configured service, and speech generation runs in the agent backend. Review the deployment and connected services before sharing sensitive audio.
Let the spoken thought become useful work.
Bring speech into the same place as your notes, tasks, and research. Ghost is in private beta for macOS.
Request access