Skip to main content
Docs Flashcards

Audio and text-to-speech

How flangu generates and plays spoken audio for flashcard words and example sentences: where audio appears, how playback works, caching, Pro requirements, and how audio differs from pronunciation text.

On this page

Every card panel has a speaker icon speaker icon. Tapping it plays a spoken pronunciation of the word or phrase on that side. Each example sentence can also be played. This article explains how audio is generated, when it is available, and how it differs from the written pronunciation field.

Where audio appears

A card has four distinct pieces of audio.

Audio What it plays
Front panel The word or phrase on the front of the card
Back panel The translation on the back of the card
Front example sentences Each sentence in the front panel's examples icon Example sentences section
Back example sentences Each sentence in the back panel's examples icon Example sentences section

Each piece of audio is independent. Tapping the front speaker icon plays only the front word. Tapping an example sentence plays only that sentence.

How to play audio

To play the word on a card panel: tap the speaker icon speaker icon in the icon row at the bottom of that panel.

To play an example sentence: open the examples icon Example sentences section, then tap the sentence text.

The speaker icon fills while audio is playing and returns to its default state when playback finishes.

On-demand generation and caching

Audio is not pre-generated for every card. The first time you tap the speaker icon for a word or phrase, flangu generates the audio and plays it. That audio is then saved locally on your device.

Every subsequent tap plays instantly from the local cache. No network call is made. Once audio has been generated for a word, it is available offline from that point on.

The same pattern applies to example sentences. Each sentence's audio is generated once and saved at its position in the list.

Note

If another user has already generated audio for the exact same text and language combination, the existing file is reused automatically. In that case, the first tap is also instant.

Voice and language

The voice used for playback is determined by the language assigned to that card panel. The front panel uses the front language's voice; the back panel uses the back language's voice. Example sentences on each side follow the same voice as that panel.

The voice cannot be changed per-card. It is set by the language configuration and applies uniformly to all cards using that language.

Pro and offline behavior

Note

High-quality cloud audio requires flangu Pro and an internet connection.

flangu Pro, online: tapping the speaker icon generates and plays a high-quality synthesized voice matched to the card's language.

Free tier or offline: the app uses the device's built-in speech engine instead. Playback is immediate but the voice quality and accent may differ from the cloud version. This fallback is automatic. No action is required.

Cloud audio also falls back to the device's built-in engine if the request fails for any reason.

Audio versus pronunciation text

The Pronunciation field and the speaker icon speaker icon are separate features that serve different purposes.

Pronunciation is text: a written romanization or IPA transcription of how the word sounds. It appears as a section in the icon row, the same as definitions or synonyms.

The speaker icon plays synthesized speech: an audio recording of the word or phrase spoken aloud.

Both can appear on the same card at the same time. The pronunciation field lets you read how the word sounds; the speaker icon lets you hear it.