Skip to content
/a/ FaceCue Performance Studio

Convert Speech

An actor's read moved onto another voice. The performance stays yours, the voice changes.

Synthesis makes a line from text and decides the delivery for you. Conversion is the opposite trade: somebody performs the line, with the timing and emphasis they meant, and the conversion keeps all of that while changing who it sounds like.

It is the tool for a small cast covering a large one, for a placeholder read by whoever was in the room becoming a character's voice, and for keeping one actor's performance when the character they voiced has been recast.

It Keeps the Performance, Not the Voice

Everything about how the line was said carries across: the timing, the stress, the pauses, the rise and fall. What changes is the timbre.

So a flat read converts into a flat read in a different voice. If the delivery is wrong, converting it will not fix it.

Input

Single takes one clip, an AudioClip from your project or a path to a file. Batch takes a folder and skips clips that already have a converted version, so a re-run only does what is new.

Each converted clip is named after its source, which keeps a batch traceable and keeps the naming that pairs a clip with its cues.

Voice

The target is whichever voice is selected in Registered Voices. Building one is covered in Voices.

Cleaning the Input

Conversion reproduces what it is given, including the room it was recorded in. Denoise cleans the source first, which matters more here than it would elsewhere, because noise that survives into the conversion is noise the new voice appears to have made.

It is worth turning on for a recording made on a poor microphone or in a live room, and worth leaving off for studio audio, where it has nothing to remove and can only take something away.

Two controls shape it. A threshold decides how aggressive the cleaning is, where higher cleans almost everything and takes some of the voice with it. A bypass skips the cleaning entirely on clips that are already clean enough to leave alone, which is what you want across a batch of mixed quality.

Denoise is a separate download. The window offers it the first time you reach for it.

Polishing the Output

Conversion reproduces the source's tone, which can leave the result sounding slightly duller than the original. Two optional lifts address that, and both are gentle by design.

Presence adds detail at the top end that the conversion itself cannot produce. Higher stays right at the top and is heard as air around the voice.

Brightness lifts the output more broadly. Higher is brighter, and pushed too far it starts to hiss.

Both blend against the raw output instead of replacing it, so turning the amount to nothing leaves the conversion untouched. Start with them off. They are corrective, and a conversion that already sounds right does not need them.

Output

The converted clip is written to the output folder, named from its source. The same overwrite choices apply as elsewhere in the window: keep what is there, replace it, or be asked.

Then Bake It

A converted clip is an ordinary recording as far as the rest of FaceCue is concerned. Bake it, register it, and it plays like anything else.

Bake the converted clip, not the original. The two have the same timing, so reusing the original's bake looks like a shortcut. Bake the one that ends up in your game: cues describe the audio they were made from, and that is the converted take.