Search results for "speech"
Found 45 results (43 tools · 1 articles · 1 skills). Sorted by relevance by SeoAIu.
AI Tools (43)
IFlytek Smart Creation: AIGC, a one-stop intelligent dubbing and virtual human creation platform
246IFlytek Zhizuo is an AIGC content creation platform launched by iFlytek, officially released in April 2023, positioned as a one-stop intelligent dubbing and virtual human application service platform. The platform utilizes AI core technologies such as speech recognition, semantic understanding,
2026-06-02Adobe Podcast: The free AI tool that transforms smartphone recordings into studio-quality audio.
148Adobe Podcast is a free AI-powered audio enhancement and podcast production tool from Adobe; its "Enhance Speech" feature can instantly elevate smartphone recordings to studio quality. It supports remote recording and text-based editing, and the free version is already highly capable.
2026-06-30A Quick Take: Just how much hassle can that AI meeting note-taker hidden inside Baidu Netdisk really save you?
145"Jiandan Tingji" is an AI-powered speech-to-text tool launched by Baidu Netdisk. Integrated with the ERNIE Bot (Wenxin Yiyan) large language model, it supports the transcription of meetings, interviews, and lectures, as well as the intelligent generation of meeting minutes, achieving an accuracy rate of up to 97%. It is available across all platforms, with a continuous monthly subscription priced at 25 yuan.
2026-06-29YiQiJian - Free AI Video Editing & Intelligent Creation Tool
129YiQiJian is a powerful free online AI video editing tool that offers a massive library of materials, exquisite video templates, intelligent text recognition, text segmentation, text-to-subtitle, speech-to-subtitle, smart voice dubbing, and automatic matching of materials and templates. It supports cloud-based automatic video synthesis and one-click publishing to mainstream video platforms, helping self-media and content creators produce videos at zero cost and achieve rapid multi-channel distribution. Key features include AI digital humans, real-time trending news, smart video planning, deep content creation, and AI auto-editing, making it the world's first free AI video creation agent.
2026-07-17iFlytek AI University – iFlytek's official AI learning and training platform, providing one-stop coverage from model fine-tuning to application development.
123iFlytek AI University is an AI professional learning and training platform under iFlytek's open platform. Leveraging iFlytek's technological expertise in NLP, intelligent speech, and other fields, it offers a full-chain course from beginner to advanced model fine-tuning and agent development. The platform is deeply integrated with iFlytek's MaaS tools, supporting zero-code model customization and deployment. Upon completion, students receive official certification and it serves as a core hub for offline developer communities in multiple cities across China.
2026-07-14Vocu AI — A domestic large-scale speech model ranked #1 globally on the HuggingFace TTS Arena, capable of infusing AI-generated speech with a human touch and genuine emotion.
120Vocu AI is an AI speech synthesis platform developed in-house by Guangzhou Shuogu Technology; its V3 series models ranked first globally in the Hugging Face TTS Arena blind tests. It offers both instant cloning (requiring a 3-second sample with over 95% similarity) and professional-grade cloning. Supporting more than 30 languages and dialects, the platform delivers precise emotional expression and cinematic-quality performance. API access is available to facilitate efficient integration for developers.
2026-07-11Tongyi Listening and Comprehension – No more frantically typing on the keyboard during meetings and lectures, a free AI tool that automatically helps you take notes.
118Tongyi Listening is an AI audio and video assistant launched by Alibaba Cloud. Based on the Tongyi big data model, it enables real-time speech-to-text conversion, intelligent summarization, multilingual translation, and speaker identification. It supports uploading audio and video files up to 6 hours long and generates meeting minutes, mind maps, and to-do lists with one click. The free version provides 48 hours of real-time recording credits daily. It is suitable for professionals, students, journalists, and other groups who frequently need to process audio and video content.
2026-07-02ElevenLabs—a powerful voice synthesis tool that makes AI speak with the same emotion as a real person, handling audiobooks, podcasts, and games all in one.
110ElevenLabs is an AI audio research company that provides products such as text-to-speech, voice cloning, and voice agents. Its AI models can generate human voices with natural intonation, emotion, and contextual understanding, supporting speech recognition for over 70 languages and over 90 other languages. The platform offers over 10,000 voice options and is trusted by over 7.5 million creators and businesses, widely used in audiobooks, podcasts, video games, customer service, and other fields.
2026-07-02Deepgram—a leading enterprise-grade voice AI platform offering real-time speech recognition, synthesis, and fully managed voice agent APIs.
95Deepgram is a leading voice AI platform that provides developers with high-accuracy, cost-effective real-time speech-to-text (STT), text-to-speech (TTS), and a unified voice agent API. Its Nova series models outperform competitors in both accuracy and speed; Aura-2 TTS offers latency under 200 milliseconds, and the voice agent API is priced at just $4.50 per hour. Supporting both cloud and self-hosted deployments, the platform is trusted by over 200,000 developers.
2026-07-11FakeYou—An AI celebrity voice and video generator that lets "anyone" say whatever you want to hear.
94FakeYou is an AI-powered voice and video generation tool that allows users to generate audio from text or speech using a vast library of voices, including those of celebrities and anime characters. It supports basic voice cloning—with some voices trained by the community (unofficial)—and is suitable for entertainment, meme creation, and creative content.
2026-07-11Ciniaoniao Voiceover—a permanently free AI voiceover tool; over 200 voices transform text into expressive speech in seconds.
92Ci Niao Voiceover is a free, AI-powered text-to-speech application featuring over 200 voice options and support for Mandarin, Cantonese, English, and various dialects. It offers nearly 300 distinct voices and more than ten emotional styles, catering to use cases such as short-video voiceovers, film and TV commentary, and audiobooks. The service is accessible across multiple platforms, including the web, mobile apps, and mini-programs.
2026-07-11MiniMax Audio—an ultra-realistic large-scale speech model capable of everything from 10-second voice cloning to support for over 40 languages, enabling AI to speak with the warmth of a real person.
89MiniMax Audio is an AI voice platform under MiniMax that offers features such as text-to-speech, voice cloning, and music generation. Its proprietary Speech series models support over 40 languages, ultra-long text, emotional expression, and low-latency interaction, and have been adopted by leading global platforms and products such as LiveKit, Pipecat, Gaotu, and Ximalaya.
2026-07-11ModelScope – China's largest and most active open-source AI model platform, boasting a "model-as-a-service" ecosystem with over 30 million developers.
88ModelScope is a leading open-source AI model platform in China, initiated by Alibaba DAMO Academy in conjunction with the CCF Open Source Development Committee. Adhering to the "Model as a Service" (MaaS) concept, the community has gathered over 200,000 open-source models, 30 million developers, and over 500 contributing organizations. It provides end-to-end services from model experience, download, fine-tuning, training to deployment, covering LLM, multimodal, speech, and AIGC fields. As a core infrastructure for AI developers in China, it is promoting the democratization of AI technology through "open source."
2026-07-14TTSMaker—A free online text-to-speech tool; generate over 200 voice styles with a single click.
88TTSMaker is a free online AI text-to-speech tool that supports multiple languages and over 200 voice styles. No registration is required; simply enter text to generate natural, realistic speech, adjust speed and volume, and download MP3 files. Ideal for video voiceovers, audiobook production, and educational content creation, it allows text to easily "come to life" with a voice.
2026-07-11Uberduck—an all-in-one voice studio where AI speaks, sings, and raps, with over 5,000 voices at your disposal.
87Uberduck is an AI-powered platform for voice and music creation, offering features such as text-to-speech, voice cloning, text-to-singing/rapping, and AI music generation. With a library of over 5,000 expressive voices and support for more than 70 languages and hundreds of musical styles, the platform enables creators, musicians, and marketers to rapidly produce professional-grade audio content.
2026-07-11LOVO AI — A lifelike TTS platform with an integrated AI video editor, featuring over 500 voices and more than 100 languages.
87LOVO AI is a high-fidelity text-to-speech and video creation platform. Its core product, Genny, integrates voice generation with online video editing, offering over 500 voices and 100 languages, alongside features such as AI scriptwriting and automatic subtitling. Supporting voice cloning and team collaboration, the platform is ideal for content creators, marketers, and educators.
2026-07-11Langlang Voiceover—a permanently free AI voiceover tool featuring over 1,100 voice talents, support for 80+ languages, and nearly 20 fine-tuning options.
86Langlang Voiceover is a permanently free AI voiceover platform featuring over 1,100 AI voices, support for more than 80 languages, and over 10 emotional styles. It offers nearly 20 fine-tuning features—such as continuous reading, pauses, handling of polyphones (characters with multiple pronunciations), localized speed adjustment, and multi-speaker narration—allowing for word-by-word customization of the voice output. The platform includes a built-in library of royalty-free (CC0) background music and provides productivity tools for subtitle generation, batch processing, and text extraction, making it ideal for short-video voiceovers, audiobook production, advertising, and more.
2026-07-11Noiz AI—More than just "voice cloning": recreating the "physicality" of the digital world through voice models.
85Noiz AI is an AI technology company specializing in comprehensive audio generation, covering speech, sound effects, ambient sounds, and music. Its AudioX-Turbo model supports "Anything-to-Audio" capabilities, enabling the generation of 10 seconds of high-quality audio within 0.24 seconds from text, video, or image inputs; the model has been open-sourced and serves approximately 1.2 million users worldwide.
2026-07-11AssemblyAI—the speech recognition API of choice for developers—build next-generation speech AI applications with industry-leading accuracy.
83AssemblyAI offers a production-grade speech-to-text API that supports 99 languages with over 93.3% word accuracy and pricing as low as $0.15 per hour. The platform integrates speech understanding capabilities—such as speaker diarization, PII identification, and summarization—and supports real-time streaming transcription with latency as low as 300ms; it is used by over 200,000 developers to build applications such as voice agents, meeting recorders, and conversation analytics tools.
2026-07-11NaturalReader—an AI text-to-speech tool chosen by over 10 million users, suitable for everything from personal reading to commercial voiceovers.
83NaturalReader is an AI text-to-speech tool that offers natural-sounding voice narration and supports various formats, including PDFs, web pages, and images. It features AI podcast generation, voice cloning, and multilingual voiceovers, making it suitable for personal learning, educational settings, and commercial video narration. It has served over 10 million users worldwide.
2026-07-10Memo AI — A locally running AI tool for transcribing audio and video to text.
82Memo AI is an all-in-one, local audio-to-text tool that effortlessly converts content from YouTube, podcasts, and local files into transcripts. It offers features such as transcription and translation for over 90 languages, AI-generated summaries, AI mind maps, and text-to-speech capabilities. Operating entirely locally to ensure privacy, it supports Windows and macOS and leverages GPU acceleration for M-series chips, enabling a 30-minute audio file to be transcribed in just two minutes.
2026-07-11Clipchamp—Microsoft's official AI video editor: edit videos for free in your browser.
80Clipchamp is an online video editor from Microsoft that integrates features such as AI-powered auto-captions, text-to-speech, and noise suppression, while also including a built-in library of royalty-free stock assets. Requiring no downloads and supporting high-quality exports without usage limits, it offers a one-stop solution—covering everything from recording to editing—for Windows and Mac users.
2026-07-10Tingnao AI—a powerful speech-to-text tool with 98.7% accuracy; generate minutes for a one-hour meeting in just three minutes.
79Tingnao AI is an AI-powered tool for audio-to-text transcription and meeting minutes generation. It supports real-time transcription, file uploads, and the processing of online audio and video content. Powered by advanced models like DeepSeek-R1, it features automatic speaker identification, intelligent summarization, and action item extraction, enabling the processing of a one-hour meeting in just three minutes. It supports multiple languages—including Chinese, English, Japanese, and Korean—and is available for an annual fee of just 199 RMB, making it ideal for professionals, students, and content creators who frequently handle audio and video materials.
2026-07-11WellSaid Labs—Give your content a "premium" voice using AI speech licensed from professional voice actors.
78WellSaid Labs is an enterprise-grade AI voice synthesis platform; its AI model, Caruso, is trained exclusively on audio licensed from professional voice actors. The platform delivers hyper-realistic, commercially viable AI voices—supporting fine-grained, word-by-word tuning as well as multiple languages and dialects—and is used by over half of the Fortune 500 companies, as well as organizations such as NPR and LinkedIn.
2026-07-11MetaVoice—A "full-duplex" voice AI platform that makes AI voice conversations as natural and expressive as those between real people.
78MetaVoice is a company specializing in the development of "Duplex" voice AI, aiming to imbue AI conversations with human-like naturalness and emotional intelligence. Its core model, MetaVoice-1B, features 1.2 billion parameters and supports emotionally expressive speech synthesis and few-shot voice cloning, dedicated to delivering immersive conversational experiences for applications such as sales, psychological counseling, and gaming.
2026-07-11TTSMaker — A free, commercially usable AI text-to-speech tool offering a choice of over 300 voices across more than 50 languages.
78TTSMaker is a free, no-registration AI text-to-speech tool that supports over 50 languages and more than 300 voice styles. You can convert text into natural-sounding speech in just three steps without signing up, and the generated audio can be used for commercial purposes—such as video voiceovers and audiobook production—free of charge. It offers a free monthly allowance of 30,000 characters, with unlimited usage available for select voices.
2026-07-11iFLYTEK Tingjian—More than just transcription: an AI-powered voice recording and insight partner that truly understands you.
78iFLYTEK Tingjian is an AI-powered voice recording and transcription platform under iFLYTEK. It offers services such as real-time speech-to-text, transcription of audio and video files, AI-driven summarization, and adaptive meeting minute generation. Powered by large models like iFLYTEK Spark, the platform supports mixed-language recognition (Mandarin, English, Cantonese) and various regional dialects; it can transcribe one hour of audio in as little as five minutes with an accuracy rate of up to 98%. Catering to both individual and enterprise users, it provides a one-stop solution that spans everything from meeting recording to the accumulation of knowledge assets.
2026-07-11WellSaid Labs—Give your content a "premium" voice using AI speech licensed from professional voice actors.
76WellSaid Labs is an enterprise-grade AI voice synthesis platform; its AI model, Caruso, is trained exclusively on audio licensed from professional voice actors. The platform delivers hyper-realistic, commercially viable AI voices—supporting fine-grained, word-by-word tuning as well as multiple languages and dialects—and is used by over half of the Fortune 500 companies, as well as organizations such as NPR and LinkedIn.
2026-07-11iFLYTEK Dubbing — A one-stop creation platform for AI-powered intelligent dubbing and virtual digital humans, driven by the iFLYTEK Spark Large Model.
76iFLYTEK Dubbing is a professional AI voiceover and virtual digital human creation platform launched by iFLYTEK. Powered by the iFLYTEK Spark large model and advanced speech synthesis technology, it offers over 1,000 hyper-realistic voices, emotional tone control, one-shot voice cloning, and the ability to generate virtual digital human videos in just two steps. The platform supports mixed Chinese-English audio, 12 dialects, and multiple foreign languages; it is suitable for use cases such as short videos, promotional films, and educational training, and has already served tens of thousands of users.
2026-07-11Voicemaker — A cost-effective AI voiceover platform featuring over 130 languages and 1,500+ voices, with unlimited use of free standard TTS.
76Voicemaker is an AI text-to-speech tool offering a selection of over 130 languages and 1,500+ voices, with support for emotion control, voice cloning, voice-to-voice conversion, and API integration. Its ProPlus models support SSML and voice effects, while the free version offers unlimited standard voiceovers. It is suitable for video narration, course creation, and IVR systems.
2026-07-11There are 13 more tools not shown. Try more specific keywords.