Skip to content
Back to course

4. Voice, audio and music with AI

Welcome to Lesson 4! Until now you have worked with text and images. This lesson introduces the world of AI-generated audio — realistic voice-overs, sound effects, podcast intros, background music, and even full songs. By the end you will know which tools to use, how to prompt them, and the ethical rules that keep your brand safe.

For decades, producing professional audio required expensive equipment, a recording studio, and trained voice actors or musicians. A 30-second radio ad jingle in Addis Ababa could cost thousands of birr by the time you paid the studio, the sound engineer, and the performer.

AI audio generation changes that equation completely. Today you can:

• Type a script in Amharic or English and receive a natural-sounding spoken voice-over in seconds — no recording booth needed.

• Describe a musical mood in one sentence and receive an original background music track — no musician needed.

• Generate a sound effect (a coffee-cup clink, a market crowd noise, a phone notification) from a short text description.

• Clone an existing voice with just a few minutes of sample audio — useful if a brand spokesperson has a distinctive voice and cannot be in the studio every week.

For small businesses, content creators, and startups across Ethiopia, this means professional-quality audio is now within reach at near-zero cost.

promptAI
AI audio tools take a text prompt or sample audio as input and generate a voice-over, music track, or sound effect as output.

AI audio tools fall into three main branches. Understanding each one helps you pick the right tool for the right job.

── BRANCH 1: AI VOICE-OVER (Text-to-Speech) ──

You paste a script; the AI reads it in a chosen voice, language, accent, and pace. Modern tools sound strikingly human — not the robotic voices of old GPS devices.

Leading tools:

• ElevenLabs — the industry benchmark for natural-sounding voices; supports many languages.

• Murf AI — business-focused; good for presentations, e-learning, and ads.

• Google Text-to-Speech / Amazon Polly — reliable, affordable, API-friendly for developers.

• Descript — lets you edit audio as if it were text; delete a word in the transcript and it disappears from the audio.

── BRANCH 2: AI MUSIC GENERATION ──

You describe the style, mood, instruments, or lyrics; the AI composes and produces a complete music track.

Leading tools:

• Suno — type a description like "upbeat Ethiopian-inspired market music, drums and masenqo, 60 seconds" and receive a complete track.

• Udio — similar to Suno; strong on diverse world music styles.

• Stable Audio (Stability AI) — precise control over length and style; good for professional producers.

• Soundraw — designed for content creators; royalty-free output.

── BRANCH 3: AI SOUND EFFECTS ──

You describe a sound; the AI synthesises it from scratch.

Leading tools:

• ElevenLabs Sound Effects — type "crowd cheering at a football match in Addis Ababa" and receive a realistic WAV file.

• Adobe Firefly Audio (in development) — integrated with Adobe's creative suite.

• Freesound + AI search — find and combine community sounds with AI-assisted filtering.

All the tools above have free tiers you can try right now. ElevenLabs free tier gives you 10,000 characters of voice synthesis per month — enough to voice a 5-minute audio ad. Suno free tier gives you 50 songs per day. Start experimenting without spending any birr.

Just as a well-structured prompt produces a better image, a well-structured audio prompt produces better sound. The exact format differs slightly for voice-overs versus music.

FOR VOICE-OVER PROMPTS (what you put in the script field and settings):

• Write the script naturally, as if a human will read it — punctuation, commas, and sentence length affect pacing.

• Choose the right voice: most tools let you pick from dozens of AI voices by gender, age, and accent. Pick one that matches your brand tone.

• Set the speaking rate: slower for formal announcements, faster for upbeat ads.

• Specify emotion if supported: "warm and friendly", "confident and authoritative", "excited".

Example script for an Abebe Coffee ad:

«ቡናዎን ፈልጎ ጠፋ? አቤቤ ቡና ዛሬ ደጃፍዎ ያደርሳል። ትዕዛዝ ለመስጠት ቴሌብርር ላይ 8796 ይደውሉ — ነጻ ዴሊቨሪ ለመጀመሪያ ትዕዛዝ!»

[Settings: Female voice, Warm tone, Medium pace]

FOR MUSIC GENERATION PROMPTS:

A good music prompt covers:

1. MOOD / EMOTION — e.g. "joyful", "tense", "peaceful", "energetic"

2. GENRE / STYLE — e.g. "Afrobeats", "Ethiopian traditional", "lo-fi hip-hop", "corporate"

3. INSTRUMENTATION — e.g. "masenqo, drums, soft piano", "acoustic guitar only"

4. DURATION — e.g. "60 seconds", "30-second loop"

5. TEMPO — e.g. "slow at 70 BPM", "upbeat at 120 BPM"

Example for Suno:

"Joyful Ethiopian market scene, traditional krar and drums, upbeat at 110 BPM, 45 seconds, no vocals, suitable as a background track for a short promotional video in Addis Ababa."

That level of detail consistently produces usable, professional-sounding results.

Here is a practical workflow that Almaz, a catering business owner in Addis Ababa, could use to create a 30-second Amharic radio-style audio ad — entirely with AI, in under 20 minutes, for zero extra cost.

STEP 1 — WRITE THE SCRIPT WITH AI

Open ChatGPT or Claude and prompt: "Write a 30-second radio ad script in Amharic for Almaz Catering, specialising in traditional Ethiopian food for weddings and corporate events in Addis Ababa. Tone: warm and professional. Include a TeleBirr payment reminder at the end."

Review and adjust the script.

STEP 2 — GENERATE THE VOICE-OVER

Open ElevenLabs (free tier). Paste the Amharic script. Choose a warm female voice. Set the pace to medium-slow for clarity. Click Generate. Download the MP3.

STEP 3 — GENERATE BACKGROUND MUSIC

Open Suno (free tier). Prompt: "Gentle, warm Ethiopian music, soft krar and flute, 30 seconds, no vocals, suitable background for a catering radio ad." Download the best track.

STEP 4 — MIX THE AUDIO

Open a free audio editor such as Audacity (free, downloadable) or Adobe Podcast (free online). Import the voice-over and the music track. Lower the music volume so it sits behind the voice. Export as MP3.

STEP 5 — REVIEW AND PUBLISH

Listen carefully — check that every word is clear, that the music does not overpower the voice, and that the script is accurate. Get a colleague to listen. Then post to social media, share on WhatsApp, or deliver to a radio station.

Total cost: 0 birr (free tiers). Total time: approximately 15–20 minutes.

Real-world result: A small school-supplies shop in Piassa generated a 45-second Amharic back-to-school jingle using Suno (music) + ElevenLabs (voice-over) + Audacity (mix). The owner shared it on five neighbourhood WhatsApp groups. Sales increased by 35% in the following week compared to the same week the year before — a result that previously required hiring a full radio-production team.

AI audio raises unique ethical and legal risks that do not apply to images or text in the same way.

1. VOICE DEEPFAKES — using AI to make someone's voice say something they never said is one of the most dangerous misuses of generative AI. It has been used to spread political disinformation, commit fraud (scammers have called families pretending to be a kidnapped relative asking for TeleBirr transfers), and damage reputations. This is illegal in most jurisdictions and deeply unethical.

2. MUSIC COPYRIGHT — AI music generators are trained on existing music. Some tools may produce output that sounds very similar to a specific copyrighted song. Before using AI music commercially:

• Use tools with royalty-free output agreements (Soundraw, Suno's commercial plan, Udio's commercial plan).

• Run the output through a copyright detection service like TuneCore's cover-song tool or SoundChain before publishing on platforms like YouTube or TikTok.

• If you are unsure, commission a human composer or use a free-licence music library (e.g. YouTube Audio Library).

3. CONSENT FOR REAL VOICES — never use a recording of a real person's voice without their consent, whether for cloning or as training data.

4. TRANSPARENCY — if your brand uses an AI voice for customer service calls, disclose this to customers. In some countries (and increasingly in Ethiopian draft digital regulations) failing to disclose AI use in commercial audio is a regulatory offence.

5. MISREPRESENTATION — do not use AI to create fake testimonials, fake radio interviews, or fake news clips in any voice. This is fraud.

Voice fraud is rising in Ethiopia. Scammers use voice-cloning apps to impersonate relatives, bosses, or bank staff and demand urgent TeleBirr transfers. If you receive a surprising voice message from someone asking for money — even if it sounds exactly like them — always verify through a separate call or in-person before transferring anything. And never use AI voice tools yourself for deception.

Scenario

Abebe runs a small FM radio programme in Bahir Dar. He wants to use AI to create a 20-second jingle promoting his Saturday morning show. He generates a track with Suno, then records his own voice introducing the show with ElevenLabs as an AI-enhanced version of his voice. He mixes them in Audacity. What should he do BEFORE airing this on his FM broadcast?

Lesson recap — Voice, audio and music with AI: • AI audio tools fall into three branches: voice-over (text-to-speech), music generation, and sound effects. • Key voice tools: ElevenLabs, Murf AI, Descript. Key music tools: Suno, Udio, Soundraw. All have free tiers. • Write effective prompts: for voice-over — script + voice style + pace + emotion; for music — mood + genre + instruments + duration + tempo. • Voice cloning is legitimate with explicit written consent; without consent it is fraud. • Music copyright: always verify your tool's commercial licence before broadcasting or publishing on platforms. • Practical workflow: AI script → AI voice-over → AI music → mix in Audacity → review → publish. Total cost: 0 birr on free tiers. • Never use AI audio to impersonate, deceive, or create false testimonials. • Next: Lesson 5 will cover AI video generation.

Check your understanding

1/7 · 82 XP

Which of the following best describes what an AI text-to-speech (voice-over) tool does?