Gadgets & Reviews

No Mic Needed: You Can Create Music and Speech With Adobe’s AI Audio Tools

[post_content]


Disclaimer: This article has been automatically aggregated from

Adobe announced Thursday that it’s adding to its vast collection of artificial intelligence tools, this time with a focus on audio. The new Firefly AI tools use AI to generate speech, music and sound effects, forming “a very stable foundation” of audio options for creators, said Jay LeBoeuf, Adobe’s head of AI audio.

Unlike Suno or other AI music generators that can create entire songs in seconds, Adobe is offering more targeted, professional-grade tools for filmmakers, musicians and creators. Think about AI converting creators’ scripts into audio to be overlaid on a TikTok video or creating custom soundtracks without worrying about copyright infringement, thanks to Adobe’s universal license.

“We’re not trying to be somebody’s wedding music here,” LeBoeuf said in an interview. The goal is to build AI “tools that are useful” and address pain points in the audio creation and editing process.

To use the new tools, you’ll need access to Firefly, Adobe’s AI hub, which may be included in your current Creative Cloud subscription, depending on your specific plan or your company’s AI permissions. AI audio generations will count as generative credits, so keep an eye on how quickly you use those up. You can nab a Firefly-only subscription starting at $10 per month.

AI audio that isn’t robotic

To create speech, upload a script you’ve written, and it will transform it into an audio file. You’ll be able to choose from several artificial voices in a variety of genders and ages, and you can translate the audio into over 20 languages.

If you have names or products that are hard for the AI to read, you can add pronunciation guidance. My last name, Chedraoui, for example, could be phonetically spelled out as “Shed-rao-wee,” instead of whatever hideous sound the AI produces when pronouncing four vowels in a row.

One of the biggest challenges with AI audio is making voices sound less robotic. Monotonous audio is boring to listen to, and it’s a clear sign of AI. Adobe built its AI audio model to recognize and apply various emotions to its outputs. When you use Adobe’s generate speech tool you can use “emotion tags” to direct the AI to apply different expressions.

To help the models understand emotions, Adobe collected more emotive training material, LeBoeuf said. “Our design team knew that we were going to control them with these adjectives and these verbs. So because it’s been part of the training since the get-go, we have this nice vertically integrated stack that allows for the highest quality expressiveness.”

Adobe’s AI policy says that it only uses licensed and publicly available content and data to train its AI models. The company says it never trains on customers’ work to improve its services.

Copyright-friendly AI music and futuristic sound effects

When generating music or soundtracks, the tools are primarily meant to create background audio — the Firefly tunes are instrumental only — not full AI songs (which you likely won’t want to listen to anyway).

You can create this background music using a Mad Libs-style fill-in-the-blank format. Choose the vibe, genre and situation you want your AI music to reflect, like a “dreamline song, with electronic, ambient style, for a video game.”

An example of generating music in Adobe Firefly, reading: A dreamlike song, with electronic and ambient style for a video game.
You can add as much as or as little detail to your AI music prompts as you want.Adobe

If you’re not sure where to start, you can let Firefly do the scoring for you. Upload a video, and the AI can write a prompt and create four sample audio tracks — up to 30 seconds long — that might fit your video’s vibe.

Sound effects make up the final part of the AI audio triad and require a prompt or an uploaded recording. For example, you can upload a video where you try to create the sound effect you want. The AI will take your human voice and transform it into whatever you want, like deepening your roar to sound like a dinosaur or monster, syncing it to the video.

Most importantly, all audio created by Adobe’s AI is commercially safe, and the music is automatically granted a universal license. That’s critical for musicians and video creators, the majority of whom have run into licensing issues before, according to a recent survey from Berklee College of Music. Social media platforms can penalize users for sharing videos that contain copyrighted music without permission. The universal license means you don’t have to worry about being personally sued for copyright infringement.

‘Control and agency’

Adobe’s new AI audio tools were first teased at last year’s Adobe Max conference, where they were released in beta. Now generally available, they add to the path that Adobe has been hurtling down to integrate generative AI into every one of its flagship programs.

And AI is already part of the daily work, said Mark Ethier, executive director of the Berklee Emerging Artistic Technology Lab. One in five surveyed musicians and creators use generative AI at some point in their process, and one-third use AI-generated audio in their final cuts.

But AI audio, specifically music, is not immune to the controversies and issues that surround AI images and video. We struggle to hear with our naked ears, if you will, the difference between AI and human-created music. After strenuous push from listeners, streaming platforms like Spotify and Deezer are adding labels to AI-generated music. Suno is also adding labels, following backlash after users learned the service trained its models on YouTube videos.

“The most important thing we heard is the desire for creators to have expressive control and agency,” Ethier said. “And really being able to have tools that support their creative process and don’t take it away.”

for informational purposes only. We do not claim ownership, accuracy, or liability for the content provided. All rights belong to the original publisher.