Voice moderation transcribes voice messages and round video notes and runs the transcript through the same anti-spam protection as text, so spam cannot hide in audio. You control whose audio is checked and how long a clip can be, with a daily transcription quota. Available on Pro and above; enable it per group in settings.
Want to set this up? Do it right in your Telm account.
Open in your dashboard1Why voice moderation
Spammers know most bots are deaf. A pitch or a scam link read aloud in a voice message or a round video note slips straight past text-only moderation.
Voice moderation closes that gap. It transcribes the audio to text and feeds that text into the very same protection that guards your group from written spam.
2How it works
When a voice message or video note arrives, the bot transcribes it and then treats the transcript exactly like a text message: rules, the ML model and, if needed, AI all weigh in.
If the spoken content is spam, it is acted on just like written spam — removed, and the author handled by your usual settings. The recovered transcript is also recorded in the moderation journal so you can see why a clip was removed.
- The audio is transcribed to text, then run through rules, the model and, if needed, AI.
- A spam transcript is acted on exactly like a spam text message.
- The transcript is logged with the decision, so nothing is a black box.
- Genuine voice notes pass through untouched — only spoken spam is acted on.
3Scope and limits
You control whose audio is checked and how long a clip can be, so checks focus where they matter most.
- Scope — check only new or untrusted members (default), or everyone in the group.
- Maximum duration — clips longer than your limit are not transcribed (default 120 seconds, adjustable up to a few minutes).
- Note: this is separate from the simple block voice messages filter, which deletes all voice notes outright without listening to them.
4Languages and accuracy
Transcription understands many languages automatically, so a pitch spoken in your community's language is transcribed and checked without any configuration.
No transcription is perfect, which is exactly why the transcript then flows through the full set of scored checks rather than triggering a punishment on its own. A muffled or ambiguous clip that scores in the grey zone goes to admin review, not an automatic ban.
- Works across many spoken languages with no manual setup.
- The transcript is scored like any text — a single mis-heard word cannot ban anyone.
- Borderline transcripts go to review rather than an automatic action.
5Turning it on
Voice moderation is a per-group setting, off by default. Choose the scope and maximum duration that fit your community.
6Plans and daily quotas
Each group has a daily limit on how many clips are transcribed, depending on your plan.
- Pro — up to 200 transcriptions per day per group.
- Business — up to 1,000 transcriptions per day per group.
- Limiting scope to new or untrusted members is the best way to stay within quota while still stopping spam.
- Overly long clips are skipped by the maximum-duration setting.
7Best practices
A little tuning gets you full coverage of the risky audio while leaving plenty of quota headroom.
- Keep the scope on new or untrusted members — trusted regulars rarely post spam by voice.
- Set a sensible maximum duration; genuine spam pitches are short, and long clips are almost never abuse.
- If you simply never want voice notes at all, use the plain block voice messages filter instead — it needs no transcription.
- Pair voice moderation with new-member restrictions so the riskiest accounts are watched most closely.