Music and audio tools

64 tools available

Back to Categories

SoundBug—a domestically developed music production software featuring built-in AI composition capabilities, removing barriers to music creation.

SoundBug is an AI-powered music production software (DAW) developed by a domestic team. It features built-in AI-driven arrangement capabilities, over 600 virtual instruments, and an extensive library of loop templates, all designed to lower the barrier to entry for music creation. Selected as a designated teaching platform for Shanghai’s "Cloud Classroom" initiative, it is highly popular among millions of young people, music teachers, and enthusiasts. Compatible with both Windows and Mac, it allows users to start creating immediately after a simple one-click installation.

Audio
53 View

Listnr—A tool featuring over 1,000 hyper-realistic AI voices and voice cloning capabilities, reaching a global audience across 142 languages.

Listnr is a powerful AI voice generation and cloning platform offering over 1,000 realistic AI voices and supporting more than 142 languages ​​and accents. Key features include rapid 30-second voice cloning, an AI video generator, and podcast hosting services, enabling content creators, marketers, and businesses to easily produce multilingual audio and video content.

Audio
50 View

Mureka—an AI music creation platform under Kunlun Tiangong—enables everyone to create professional-grade, release-ready songs.

Mureka is an AI music creation platform under Kunlun Tech, powered by the proprietary MusiCoT technology framework; its V8 and V9 models comprehensively outperform mainstream international competitors in terms of melody, vocals, and arrangement. The platform supports multi-dimensional creation—including natural language prompts, lyrics, and reference songs—and offers API access alongside AI Studio editing tools. With a monthly subscription fee of 88 RMB for the basic tier, Mureka enables AI music to evolve from mere generation to actual release.

Audio
53 View

Vocu AI — A domestic large-scale speech model ranked #1 globally on the HuggingFace TTS Arena, capable of infusing AI-generated speech with a human touch and genuine emotion.

Vocu AI is an AI speech synthesis platform developed in-house by Guangzhou Shuogu Technology; its V3 series models ranked first globally in the Hugging Face TTS Arena blind tests. It offers both instant cloning (requiring a 3-second sample with over 95% similarity) and professional-grade cloning. Supporting more than 30 languages ​​and dialects, the platform delivers precise emotional expression and cinematic-quality performance. API access is available to facilitate efficient integration for developers.

Audio
53 View

Ciniaoniao Voiceover—a permanently free AI voiceover tool; over 200 voices transform text into expressive speech in seconds.

Ci Niao Voiceover is a free, AI-powered text-to-speech application featuring over 200 voice options and support for Mandarin, Cantonese, English, and various dialects. It offers nearly 300 distinct voices and more than ten emotional styles, catering to use cases such as short-video voiceovers, film and TV commentary, and audiobooks. The service is accessible across multiple platforms, including the web, mobile apps, and mini-programs.

Audio
55 View

Kuaizhuan Subtitles—an AI-powered platform for subtitle generation and translation; a one-stop solution for multilingual video subtitles.

Kuaizhuan Subtitles is an AI-powered platform for subtitling and transcription, offering features such as audio-to-text conversion, AI subtitle generation, intelligent sentence segmentation and reformatting, and accurate translation across nearly 100 languages. It provides a one-stop solution integrating online editing, subtitle burning, and multi-format exporting (supporting MP4, SRT, ASS, etc.). With a recognition accuracy rate of up to 99%, it enables video creators, subtitle groups, and multinational enterprises to efficiently produce multilingual subtitled videos.

Audio
38 View

TTSMaker—A free online text-to-speech tool; generate over 200 voice styles with a single click.

TTSMaker is a free online AI text-to-speech tool that supports multiple languages ​​and over 200 voice styles. No registration is required; simply enter text to generate natural, realistic speech, adjust speed and volume, and download MP3 files. Ideal for video voiceovers, audiobook production, and educational content creation, it allows text to easily "come to life" with a voice.

Audio
37 View

Lemonaide AI—an AI melody generator driven by Grammy-winning producers, infusing your creations with professional inspiration.

Lemonaide AI is an AI-powered melody generation plugin driven by top-tier industry producers. It generates royalty-free MIDI melodies and chord progressions with a single click, while its exclusive "Collab Club" offers specialized AI models trained by Grammy-winning producers such as Lex Luger and Kato On The Track. Compatible with both Windows and macOS, the plugin allows for seamless drag-and-drop integration into major DAWs, making it suitable for creators ranging from beginners to professional music producers.

Audio
43 View

Langlang Voiceover—a permanently free AI voiceover tool featuring over 1,100 voice talents, support for 80+ languages, and nearly 20 fine-tuning options.

Langlang Voiceover is a permanently free AI voiceover platform featuring over 1,100 AI voices, support for more than 80 languages, and over 10 emotional styles. It offers nearly 20 fine-tuning features—such as continuous reading, pauses, handling of polyphones (characters with multiple pronunciations), localized speed adjustment, and multi-speaker narration—allowing for word-by-word customization of the voice output. The platform includes a built-in library of royalty-free (CC0) background music and provides productivity tools for subtitle generation, batch processing, and text extraction, making it ideal for short-video voiceovers, audiobook production, advertising, and more.

Audio
40 View

Tingnao AI—a powerful speech-to-text tool with 98.7% accuracy; generate minutes for a one-hour meeting in just three minutes.

Tingnao AI is an AI-powered tool for audio-to-text transcription and meeting minutes generation. It supports real-time transcription, file uploads, and the processing of online audio and video content. Powered by advanced models like DeepSeek-R1, it features automatic speaker identification, intelligent summarization, and action item extraction, enabling the processing of a one-hour meeting in just three minutes. It supports multiple languages—including Chinese, English, Japanese, and Korean—and is available for an annual fee of just 199 RMB, making it ideal for professionals, students, and content creators who frequently handle audio and video materials.

Audio
40 View

Dabing AI Voice Changer—an ultra-low-latency real-time voice changer with over 500 voices to make game streaming more fun.

Dabing AI Voice Changer is a professional real-time AI voice-changing tool featuring ultra-low latency (under 100ms) and a library of over 500 meticulously tuned voices. It requires no dedicated GPU and consumes only about 5% of CPU resources, allowing it to run smoothly even on standard laptops. Perfectly suited for scenarios such as multiplayer gaming, live streaming interactions, and social voice chats, it makes transforming your voice simple and fun.

Audio
33 View

Memo AI — A locally running AI tool for transcribing audio and video to text.

Memo AI is an all-in-one, local audio-to-text tool that effortlessly converts content from YouTube, podcasts, and local files into transcripts. It offers features such as transcription and translation for over 90 languages, AI-generated summaries, AI mind maps, and text-to-speech capabilities. Operating entirely locally to ensure privacy, it supports Windows and macOS and leverages GPU acceleration for M-series chips, enabling a 30-minute audio file to be transcribed in just two minutes.

Audio
42 View

WellSaid Labs—Give your content a "premium" voice using AI speech licensed from professional voice actors.

WellSaid Labs is an enterprise-grade AI voice synthesis platform; its AI model, Caruso, is trained exclusively on audio licensed from professional voice actors. The platform delivers hyper-realistic, commercially viable AI voices—supporting fine-grained, word-by-word tuning as well as multiple languages ​​and dialects—and is used by over half of the Fortune 500 companies, as well as organizations such as NPR and LinkedIn.

Audio
33 View

WellSaid Labs—Give your content a "premium" voice using AI speech licensed from professional voice actors.

WellSaid Labs is an enterprise-grade AI voice synthesis platform; its AI model, Caruso, is trained exclusively on audio licensed from professional voice actors. The platform delivers hyper-realistic, commercially viable AI voices—supporting fine-grained, word-by-word tuning as well as multiple languages ​​and dialects—and is used by over half of the Fortune 500 companies, as well as organizations such as NPR and LinkedIn.

Audio
35 View

OptimizerAI—Generate high-quality sound effects from text descriptions: an AI sound design workshop for game and video creators.

OptimizerAI is an AI-powered sound effect generation tool that creates custom sound effects for games, animations, videos, and more based on text prompts. It supports 44.1kHz stereo output, style customization, and the generation of audio variations, offering both free trials and paid subscriptions. It is suitable for game developers, video creators, animators, and audio designers.

Audio
39 View

Stable Audio — An AI platform from Stability AI that uses open-source models to generate high-quality, commercially viable music locally.

Stable Audio is an AI music generation platform launched by Stability AI that supports the creation of up to six minutes of 44.1kHz stereo audio from text prompts. Its open-source models can be deployed locally, the training data is fully licensed, and the generated music is explicitly cleared for commercial use. It is suitable for music producers, content creators, and developers.

Audio
39 View

Haimian Music—an AI music creation platform by ByteDance—lets you generate your own unique songs from just a single sentence or image.

Haimian Music is an AI music creation platform launched by ByteDance that enables the generation of complete songs with a single click based on inspiration, lyrics, or images. The platform offers a variety of musical styles and mood options, as well as a voice cloning feature, allowing users with no prior experience to easily create personalized music. Creations can be shared directly to social media platforms like Douyin, and the free version allows for the generation of multiple songs daily.

Audio
46 View

Deepgram—a leading enterprise-grade voice AI platform offering real-time speech recognition, synthesis, and fully managed voice agent APIs.

Deepgram is a leading voice AI platform that provides developers with high-accuracy, cost-effective real-time speech-to-text (STT), text-to-speech (TTS), and a unified voice agent API. Its Nova series models outperform competitors in both accuracy and speed; Aura-2 TTS offers latency under 200 milliseconds, and the voice agent API is priced at just $4.50 per hour. Supporting both cloud and self-hosted deployments, the platform is trusted by over 200,000 developers.

Audio
38 View

Moyin (Moyin Workshop)—an AI voiceover powerhouse featuring over 800 voices and 1,000 styles, trusted by creators of short videos and audiobooks.

Moyin (Moyin Workshop) is an AI voice synthesis platform under Mobvoi. Powered by the proprietary "Sequence Monkey" (Xulie Houzi) large model and a fifth-generation TTS engine, it offers over 800 voice profiles, more than 1,000 styles, and nearly 20 fine-tuning features. The platform supports multiple languages ​​and dialects, voice cloning, cloud-based video editing, and multi-user collaboration; it is widely used in applications such as short videos, audiobooks, and film/TV commentary, having served over 6 million users to date.

Audio
35 View

Yueyin AI Voiceover—an online smart voiceover tool from Zhipianbang, offering nearly a thousand voices for free commercial use.

Yueyin AI Voiceover is an online AI voiceover platform launched by Zhipianbang. It offers nearly a thousand voice options, including narrators with diverse styles and emotional ranges, as well as highly realistic, human-like AI voices. The platform supports fine-tuned adjustments—such as handling polyphones, pauses, and number pronunciation—and features built-in AI detection for prohibited content. Members can generate commercial usage authorizations online, making the service ideal for short videos, film and TV commentary, audiobooks, gaming, animation, and more.

Audio
36 View

MetaVoice—A "full-duplex" voice AI platform that makes AI voice conversations as natural and expressive as those between real people.

MetaVoice is a company specializing in the development of "Duplex" voice AI, aiming to imbue AI conversations with human-like naturalness and emotional intelligence. Its core model, MetaVoice-1B, features 1.2 billion parameters and supports emotionally expressive speech synthesis and few-shot voice cloning, dedicated to delivering immersive conversational experiences for applications such as sales, psychological counseling, and gaming.

Audio
31 View

Treblo—a completely free, unlimited AI music generator that turns any idea into a complete song.

Treblo (formerly Sonauto) is a completely free, unlimited AI music generation app. Users can quickly create complete songs—featuring both vocals and instrumentation—simply by providing text descriptions, lyrics, or a hummed melody. The platform offers thousands of musical styles, community sharing features, and a "Fancy Mode" for enhanced audio quality, making it ideal for content creators, music enthusiasts, and anyone interested in trying their hand at music creation.

Audio
32 View

Lyrics Into Song AI — An online music creation platform where you input lyrics and AI generates a complete song with a single click.

Lyrics Into Song AI is an online AI music generation tool that quickly transforms user-inputted lyrics into complete songs featuring melody, harmony, and arrangement. The platform offers two input modes—simple descriptions and professional lyrics—and supports a wide range of musical styles and vocal options. It also includes extended features such as an AI music editor, lyrics generator, vocal remover, and AI cover capabilities. Free users can create and preview tracks online, while paid users have access to download and commercial use rights.

Audio
39 View

Audo.ai — The ultimate one-click AI audio cleanup tool that instantly eliminates background noise from podcasts and videos.

Audo.ai (Audo Studio) is an AI-powered online audio cleanup tool that automatically removes background noise, reduces echo, and balances volume with a single click. Designed for podcasters, video creators, and remote workers, it processes audio up to 10 times faster than Adobe or Audacity. A free version offers 20 minutes of processing time per month, while paid plans start at $12 per month.

Audio
38 View

Supertone Shift — Put the voices of professional voice actors into your throat: an ultra-low-latency, real-time AI voice changer from HYBE.

Supertone Shift is a real-time AI voice-changing software developed by Supertone (a subsidiary of HYBE). It is renowned for its industry-leading low latency—as low as 47ms—and a lightweight design that operates without the need for a dedicated GPU. The software offers a vast library of AI voice personas created by professional voice actors and supports fine-tuned adjustments to parameters such as pitch, dynamic range, and reverb. Ideal for scenarios ranging from multiplayer gaming and VTuber streaming to video voiceovers, it is highly popular among creators worldwide, particularly in Japan.

Audio
36 View

Voicenotes—an AI-powered voice note and meeting recording tool that turns every conversation into a searchable knowledge asset.

Voicenotes is an AI-powered tool for voice notes and meeting records that supports one-tap recording, real-time transcription, and the automatic generation of summaries and action items. Users can ask the AI ​​questions using natural language to retrieve information from past conversations. It supports automatic recognition of over 60 languages, holds SOC 2 Type II and GDPR compliance certifications, boasts a 4.8-star rating on the App Store, and serves more than 850,000 users worldwide.

Audio
35 View

Boomy—an AI music platform that lets you create and release original songs in 30 seconds, with no musical background required.

Boomy is an AI-powered music creation platform that enables anyone to generate original songs in seconds, without requiring prior musical experience. It allows users to quickly create music by selecting a style and offers basic editing tools. Users can distribute their tracks to over 40 streaming platforms—such as Spotify and Apple Music—with a single click and earn royalties. Operating on a freemium model, the platform is ideal for content creators, podcasters, and music enthusiasts.

Audio
35 View

Beatoven.ai—an AI music generator built for creators; create emotion-driven soundtracks and ensure copyright never stands in the way of your creativity.

Beatoven.ai is an AI-powered, royalty-free music generation platform specializing in creating emotive soundtracks for videos, podcasts, games, and more. It generates unique background music from text, images, or videos, offering customization across 16 moods and various styles. Its Maestro model is trained on fully licensed data and provides royalty sharing for musicians, having already helped creators worldwide generate millions of tracks.

Audio
37 View

TTSMaker — A free, commercially usable AI text-to-speech tool offering a choice of over 300 voices across more than 50 languages.

TTSMaker is a free, no-registration AI text-to-speech tool that supports over 50 languages ​​and more than 300 voice styles. You can convert text into natural-sounding speech in just three steps without signing up, and the generated audio can be used for commercial purposes—such as video voiceovers and audiobook production—free of charge. It offers a free monthly allowance of 30,000 characters, with unlimited usage available for select voices.

Audio
30 View

Mubert—an AI platform that generates infinite royalty-free music in real-time, providing custom background music for creators, developers, and brands.

Mubert is an AI-powered, real-time music generation platform that creates unlimited, unique, and royalty-free music from text, images, or moods. It offers flexible output options ranging from 15-second soundtracks for short videos to 25-minute mixes, supports API integration and streaming, and has generated over 200 million tracks, serving 28 million creators worldwide.

Audio
41 View

iFLYTEK Dubbing — A one-stop creation platform for AI-powered intelligent dubbing and virtual digital humans, driven by the iFLYTEK Spark Large Model.

iFLYTEK Dubbing is a professional AI voiceover and virtual digital human creation platform launched by iFLYTEK. Powered by the iFLYTEK Spark large model and advanced speech synthesis technology, it offers over 1,000 hyper-realistic voices, emotional tone control, one-shot voice cloning, and the ability to generate virtual digital human videos in just two steps. The platform supports mixed Chinese-English audio, 12 dialects, and multiple foreign languages; it is suitable for use cases such as short videos, promotional films, and educational training, and has already served tens of thousands of users.

Audio
34 View

Wondercraft—an AI studio that drives video and audio creation through dialogue—cuts professional content production time from weeks to minutes.

Wondercraft is an AI-powered platform for video and audio creation. Through its built-in AI agent, "Wonda," users can produce and edit content—such as podcasts, advertisements, and training videos—using natural language conversations. The platform integrates ElevenLabs' hyper-realistic voice technology, supports over 30 languages ​​and voice cloning, and facilitates team collaboration; it has been adopted by organizations including Spotify, Amazon, and the World Bank.

Audio
42 View

Uberduck—an all-in-one voice studio where AI speaks, sings, and raps, with over 5,000 voices at your disposal.

Uberduck is an AI-powered platform for voice and music creation, offering features such as text-to-speech, voice cloning, text-to-singing/rapping, and AI music generation. With a library of over 5,000 expressive voices and support for more than 70 languages ​​and hundreds of musical styles, the platform enables creators, musicians, and marketers to rapidly produce professional-grade audio content.

Audio
44 View

Suno—an AI music powerhouse that generates full songs from a single sentence, a new creative favorite for over 100 million people worldwide.

Suno is an AI music generation platform that creates complete songs—featuring vocals, lyrics, and instrumentation—within 30 seconds based on any input idea. Version 5.5, released in March 2026, introduced features such as voice cloning, custom models, and personalized learning. The platform supports the generation of extended tracks up to eight minutes long and covers a vast array of musical styles. Free users can generate 50 songs daily, while paid plans start as low as $8 per month. Suno boasts over 100 million global users, 2 million paying subscribers, and an annual recurring revenue of $300 million.

Audio
38 View

iFLYTEK Tingjian—More than just transcription: an AI-powered voice recording and insight partner that truly understands you.

iFLYTEK Tingjian is an AI-powered voice recording and transcription platform under iFLYTEK. It offers services such as real-time speech-to-text, transcription of audio and video files, AI-driven summarization, and adaptive meeting minute generation. Powered by large models like iFLYTEK Spark, the platform supports mixed-language recognition (Mandarin, English, Cantonese) and various regional dialects; it can transcribe one hour of audio in as little as five minutes with an accuracy rate of up to 98%. Catering to both individual and enterprise users, it provides a one-stop solution that spans everything from meeting recording to the accumulation of knowledge assets.

Audio
37 View

AssemblyAI—the speech recognition API of choice for developers—build next-generation speech AI applications with industry-leading accuracy.

AssemblyAI offers a production-grade speech-to-text API that supports 99 languages ​​with over 93.3% word accuracy and pricing as low as $0.15 per hour. The platform integrates speech understanding capabilities—such as speaker diarization, PII identification, and summarization—and supports real-time streaming transcription with latency as low as 300ms; it is used by over 200,000 developers to build applications such as voice agents, meeting recorders, and conversation analytics tools.

Audio
42 View

SOUNDRAW—a 100% copyright-safe AI music generator; every track is cleared for commercial use and revenue sharing.

SOUNDRAW is an AI music generation platform that relies exclusively on a proprietary library of original tracks created by in-house producers. Users can generate unlimited royalty-free music by mixing genres and adjusting intensity and duration, with the option to export stems for professional mixing. The licensing agreement permits commercial use and distribution while allowing users to retain all royalties, making it an ideal choice for content creators and musicians.

Audio
41 View

LOVO AI — A lifelike TTS platform with an integrated AI video editor, featuring over 500 voices and more than 100 languages.

LOVO AI is a high-fidelity text-to-speech and video creation platform. Its core product, Genny, integrates voice generation with online video editing, offering over 500 voices and 100 languages, alongside features such as AI scriptwriting and automatic subtitling. Supporting voice cloning and team collaboration, the platform is ideal for content creators, marketers, and educators.

Audio
40 View

Typecast—Find the perfect voice for every story using a library of over 700 real human voices and "smart emotion" technology.

Typecast is an AI-powered platform for voice and video content creation, offering a library of over 700 voices based on real human recordings alongside voice cloning capabilities. Its core "Smart Emotion" technology automatically matches intonation and emotion to the text; the platform also supports API integration and mobile synchronization, catering to a wide range of voiceover needs—from individual creators to enterprises.

Audio
32 View

FakeYou—An AI celebrity voice and video generator that lets "anyone" say whatever you want to hear.

FakeYou is an AI-powered voice and video generation tool that allows users to generate audio from text or speech using a vast library of voices, including those of celebrities and anime characters. It supports basic voice cloning—with some voices trained by the community (unofficial)—and is suitable for entertainment, meme creation, and creative content.

Audio
37 View

Voicemaker — A cost-effective AI voiceover platform featuring over 130 languages ​​and 1,500+ voices, with unlimited use of free standard TTS.

Voicemaker is an AI text-to-speech tool offering a selection of over 130 languages ​​and 1,500+ voices, with support for emotion control, voice cloning, voice-to-voice conversion, and API integration. Its ProPlus models support SSML and voice effects, while the free version offers unlimited standard voiceovers. It is suitable for video narration, course creation, and IVR systems.

Audio
33 View

Murf AI—From voiceover studio to real-time voice agent: Generate professional-grade voices with AI.

Murf AI is an AI voice generation and conversational platform offering over 200 ultra-realistic voices, support for more than 20 languages, and voice cloning capabilities. Its Falcon TTS API enables real-time voice agent deployment with an industry-leading 55ms latency, while the Studio editor allows for audio-video synchronization, making it suitable for applications such as video voiceovers, course creation, and intelligent customer service.

Audio
37 View

Krisp—AI noise cancellation and meeting assistant for clearer, more efficient remote communication.

Krisp is an AI-powered voice platform offering two-way noise cancellation, AI accent transformation, and an intelligent meeting assistant. It provides real-time transcription, generates summaries and action items, supports multilingual translation, and integrates with meeting tools such as Zoom and Teams. Free trials are available for both individual and team plans.

Audio
42 View

Mureka — Generate complete, original songs from a single line of inspiration; supports custom vocalists and mixing.

Mureka is an AI music generation platform that allows users to create complete songs—including vocals—simply by describing their creative ideas. It offers advanced features such as custom lyrics, remixes based on reference tracks, and customizable vocal styles, all while prioritizing high-quality audio output. Paid plans start at $7.17 per month, making the platform suitable for content creators, independent musicians, and general users who value high audio quality.

Audio
39 View

MiniMax Audio—an ultra-realistic large-scale speech model capable of everything from 10-second voice cloning to support for over 40 languages, enabling AI to speak with the warmth of a real person.

MiniMax Audio is an AI voice platform under MiniMax that offers features such as text-to-speech, voice cloning, and music generation. Its proprietary Speech series models support over 40 languages, ultra-long text, emotional expression, and low-latency interaction, and have been adopted by leading global platforms and products such as LiveKit, Pipecat, Gaotu, and Ximalaya.

Audio
38 View

Noiz AI—More than just "voice cloning": recreating the "physicality" of the digital world through voice models.

Noiz AI is an AI technology company specializing in comprehensive audio generation, covering speech, sound effects, ambient sounds, and music. Its AudioX-Turbo model supports "Anything-to-Audio" capabilities, enabling the generation of 10 seconds of high-quality audio within 0.24 seconds from text, video, or image inputs; the model has been open-sourced and serves approximately 1.2 million users worldwide.

Audio
38 View

Tianpule — Create music and videos using conversational AI; generate complete works with a single click.

Tunee (Tianpule) is an AI music brand under Quwan Technology that offers an end-to-end creation platform for music and video generation. Its flagship product, "Tunee," enables the generation of complete songs—lasting up to five minutes—featuring arrangements, vocals, and visuals based on conversational prompts; meanwhile, "TemPolor Melo-D" is the world's first generative AI guitar. The platform aims to lower the barriers to music creation through a strategic approach integrating models, AI agents, and hardware.

Audio
41 View

Voice AI—a real-time voice-changing and voice-cloning tool; a "voice skin" for game streaming and content creation.

Voice AI is a real-time voice changing and cloning platform for Windows that utilizes virtual microphone technology to apply voices in real-time across scenarios such as gaming, live streaming, and meetings. It offers hundreds of voice options and custom cloning capabilities, supports integration with applications like Discord, Zoom, and OBS, and serves millions of creators and gamers worldwide.

Audio
45 View

LALAL.AI—A professional-grade AI audio track separation tool; from vocal extraction to VST plugins, it is the "audio dissector" for music producers.

LALAL.AI is an AI-powered professional audio track separation platform that utilizes deep learning models to precisely extract vocals, drums, bass, guitars, and other instruments. With the 2025 release of the Andromeda model, it ranked first among commercial tools in Meta's benchmark tests. A VST plugin supporting 7-track separation has now been launched, enabling native operation within DAWs while safeguarding the privacy of unreleased works.

Audio
42 View

Udio—an AI music creation platform that generates complete songs from text descriptions.

Udio is an AI music generation tool that allows users to create complete songs—featuring both vocals and instrumentation—simply by providing text descriptions. It offers advanced features such as audio uploading for mixing, track separation and downloading, lyric editing, and vocal cloning. Renowned for its exceptional audio quality, the tool excels particularly in musical styles driven by acoustic instruments, such as rock and jazz.

Audio
49 View

Notta—an AI meeting recording and transcription tool that turns spoken conversations into a searchable knowledge base.

Notta is an AI-powered meeting recording tool that supports real-time transcription and translation in 58 languages. It automatically generates transcripts and summaries from meetings, interviews, or lectures, and—via the "Notta Brain" feature—creates visual deliverables such as PowerPoint presentations and infographics. It is ideal for executives, sales professionals, consultants, and researchers looking to maximize the utility of information from their meetings.

Audio
41 View

Voicemod—A real-time voice changer and soundboard that brings gaming voice chat and live streams to life.

Voicemod is real-time voice-changing software featuring an extensive library of AI voices and sound effects, allowing users to instantly alter their voices or play humorous sound effects during gaming, live streaming, and voice calls. It integrates seamlessly via virtual microphone technology, enabling use without the need for additional hardware.

Audio
44 View

Flow Music — Converse with an AI music producer and create a complete song from scratch.

Flow Music is an AI music creation platform that allows users to generate complete songs by interacting with an AI "producer," supporting style control, detailed adjustments, and track separation. Integrating cutting-edge models such as Lyria 3 (for music) and Veo (for video), and offering export capabilities to mainstream DAWs, it serves as a comprehensive space for music creation—from initial inspiration to final release.

Audio
44 View

NaturalReader—an AI text-to-speech tool chosen by over 10 million users, suitable for everything from personal reading to commercial voiceovers.

NaturalReader is an AI text-to-speech tool that offers natural-sounding voice narration and supports various formats, including PDFs, web pages, and images. It features AI podcast generation, voice cloning, and multilingual voiceovers, making it suitable for personal learning, educational settings, and commercial video narration. It has served over 10 million users worldwide.

Audio
41 View

Yinjian—an all-in-one audio creation platform from Ximalaya, covering everything from noise reduction to AI-generated audiobooks.

Yinjian is an online audio editing and AI creation platform launched by Ximalaya. It offers features such as AI noise reduction, intelligent background music, AI audiobook generation, character voice switching, and automatic segmentation. Requiring no downloads and supporting simultaneous script viewing and recording, it lowers the barrier to entry for audio production, making it ideal for podcasters, audiobook creators, and audio content producers.

Audio
34 View

Clipchamp—Microsoft's official AI video editor: edit videos for free in your browser.

Clipchamp is an online video editor from Microsoft that integrates features such as AI-powered auto-captions, text-to-speech, and noise suppression, while also including a built-in library of royalty-free stock assets. Requiring no downloads and supporting high-quality exports without usage limits, it offers a one-stop solution—covering everything from recording to editing—for Windows and Mac users.

Audio
37 View

VEED.IO—An AI video studio in your browser: a one-stop solution for everything from generation to subtitles.

VEED.IO is an online AI video editing platform featuring text-to-video generation, automatic subtitles, AI voiceovers, background noise removal, and 8K upscaling. Integrating cutting-edge models such as Sora 2 and Veo 3.1, it offers a comprehensive workflow—from recording and editing to one-click publishing—trusted by teams at companies like Google and Meta, making it ideal for creators to rapidly produce high-quality video content.

Audio
35 View

IBM—A global leader in enterprise AI, hybrid cloud, and mission-critical application modernization solutions.

IBM (International Business Machines Corporation) is a globally renowned enterprise technology and consulting company that provides enterprise technologies including the watsonx AI platform, hybrid cloud infrastructure, mission-critical application modernization (such as IBM Bob™), and data and automation solutions. Its services span various industries—including IT, finance, telecommunications, and manufacturing—aiming to help enterprises achieve intelligent transformation.

Audio
39 View

Tongyi Listening and Comprehension – No more frantically typing on the keyboard during meetings and lectures, a free AI tool that automatically helps you take notes.

Tongyi Listening is an AI audio and video assistant launched by Alibaba Cloud. Based on the Tongyi big data model, it enables real-time speech-to-text conversion, intelligent summarization, multilingual translation, and speaker identification. It supports uploading audio and video files up to 6 hours long and generates meeting minutes, mind maps, and to-do lists with one click. The free version provides 48 hours of real-time recording credits daily. It is suitable for professionals, students, journalists, and other groups who frequently need to process audio and video content.

Audio
58 View

ElevenLabs—a powerful voice synthesis tool that makes AI speak with the same emotion as a real person, handling audiobooks, podcasts, and games all in one.

ElevenLabs is an AI audio research company that provides products such as text-to-speech, voice cloning, and voice agents. Its AI models can generate human voices with natural intonation, emotion, and contextual understanding, supporting speech recognition for over 70 languages ​​and over 90 other languages. The platform offers over 10,000 voice options and is trusted by over 7.5 million creators and businesses, widely used in audiobooks, podcasts, video games, customer service, and other fields.

Audio
68 View

Adobe Podcast: The free AI tool that transforms smartphone recordings into studio-quality audio.

Adobe Podcast is a free AI-powered audio enhancement and podcast production tool from Adobe; its "Enhance Speech" feature can instantly elevate smartphone recordings to studio quality. It supports remote recording and text-based editing, and the free version is already highly capable.

Audio
95 View

A Quick Take: Just how much hassle can that AI meeting note-taker hidden inside Baidu Netdisk really save you?

"Jiandan Tingji" is an AI-powered speech-to-text tool launched by Baidu Netdisk. Integrated with the ERNIE Bot (Wenxin Yiyan) large language model, it supports the transcription of meetings, interviews, and lectures, as well as the intelligent generation of meeting minutes, achieving an accuracy rate of up to 97%. It is available across all platforms, with a continuous monthly subscription priced at 25 yuan.

Audio
94 View

Yinshu AI: A Chinese speaking AI music creation community and zero threshold intelligent song

Yinshu AI is an AI music creation platform designed specifically for Chinese users, positioned as the world's first AI music community. The platform does not require music theory knowledge, and users can generate complete songs containing melodies, arrangements, and AI vocals with just one click

Audio
170 View

IFlytek Smart Creation: AIGC, a one-stop intelligent dubbing and virtual human creation platform

IFlytek Zhizuo is an AIGC content creation platform launched by iFlytek, officially released in April 2023, positioned as a one-stop intelligent dubbing and virtual human application service platform. The platform utilizes AI core technologies such as speech recognition, semantic understanding,

Audio
143 View