Search results for "voice"
Found 98 results (82 tools · 7 articles · 9 skills). Sorted by relevance by SeoAIu.
AI Tools (82)
JoyPix - AI Video Generator with AI Lip-Sync, AI Image, Avatar & Free Voice Cloning
161JoyPix.ai is an all-in-one AI creative tool featuring AI video generation, AI lip-sync, avatar creation, voice cloning, talking photo, and image generation. No camera needed—create realistic AI videos and lip-sync videos in seconds. Supports text/image to video, text/image to image, two-character dialogue, one-shot animation, and more. The Motion-2.5 series delivers precise lip-sync, physically realistic motion, and production-ready stability. Perfect for content creators, gamers, and social media users. All start for free.
2026-07-17YiQiJian - Free AI Video Editing & Intelligent Creation Tool
160YiQiJian is a powerful free online AI video editing tool that offers a massive library of materials, exquisite video templates, intelligent text recognition, text segmentation, text-to-subtitle, speech-to-subtitle, smart voice dubbing, and automatic matching of materials and templates. It supports cloud-based automatic video synthesis and one-click publishing to mainstream video platforms, helping self-media and content creators produce videos at zero cost and achieve rapid multi-channel distribution. Key features include AI digital humans, real-time trending news, smart video planning, deep content creation, and AI auto-editing, making it the world's first free AI video creation agent.
2026-07-17360AI Search – an AI search engine that can "think slowly" for you, so you no longer need to flip through dozens of pages to find information.
122360AI Search (Nano AI Search) is an AI-native search engine launched by 360, supporting multimodal input including text, voice, photos, and videos. It features a built-in "slow thinking mode," capable of calling up three or more large models at once and generating a 5,000-word in-depth report after reading 250,000 documents. It supports extracting mind maps from 100-page PDF/Word documents in one minute, extracting key points from one hour of video in one minute, and providing a one-stop AI tool for webpage AI summarization, PPT generation, and academic source tracing. In July 2024, user visits exceeded 90 million, transforming search from "providing links" to "providing direct answers."
2026-07-02ElevenLabs—a powerful voice synthesis tool that makes AI speak with the same emotion as a real person, handling audiobooks, podcasts, and games all in one.
121ElevenLabs is an AI audio research company that provides products such as text-to-speech, voice cloning, and voice agents. Its AI models can generate human voices with natural intonation, emotion, and contextual understanding, supporting speech recognition for over 70 languages and over 90 other languages. The platform offers over 10,000 voice options and is trusted by over 7.5 million creators and businesses, widely used in audiobooks, podcasts, video games, customer service, and other fields.
2026-07-02Haimian Music—an AI music creation platform by ByteDance—lets you generate your own unique songs from just a single sentence or image.
115Haimian Music is an AI music creation platform launched by ByteDance that enables the generation of complete songs with a single click based on inspiration, lyrics, or images. The platform offers a variety of musical styles and mood options, as well as a voice cloning feature, allowing users with no prior experience to easily create personalized music. Creations can be shared directly to social media platforms like Douyin, and the free version allows for the generation of multiple songs daily.
2026-07-11Rabbit R1 – An AI hardware device that challenges the "App terminator," a small orange box that went from being ridiculed to gaining popularity.
115The Rabbit R1 is an AI-native pocket assistant launched by Rabbit Inc. It features a Large Motion Model (LAM) and allows users to directly control apps and complete tasks via voice commands. Although it was controversial due to its incomplete functionality after its release in early 2024, its reputation has reversed after more than 35 OTA updates over two years and a rewrite of Rabbit OS 2. The company also plans to launch a more geeky Cyberdeck product.
2026-07-09Dabing AI Voice Changer—an ultra-low-latency real-time voice changer with over 500 voices to make game streaming more fun.
114Dabing AI Voice Changer is a professional real-time AI voice-changing tool featuring ultra-low latency (under 100ms) and a library of over 500 meticulously tuned voices. It requires no dedicated GPU and consumes only about 5% of CPU resources, allowing it to run smoothly even on standard laptops. Perfectly suited for scenarios such as multiplayer gaming, live streaming interactions, and social voice chats, it makes transforming your voice simple and fun.
2026-07-11FakeYou—An AI celebrity voice and video generator that lets "anyone" say whatever you want to hear.
112FakeYou is an AI-powered voice and video generation tool that allows users to generate audio from text or speech using a vast library of voices, including those of celebrities and anime characters. It supports basic voice cloning—with some voices trained by the community (unofficial)—and is suitable for entertainment, meme creation, and creative content.
2026-07-11Deepgram—a leading enterprise-grade voice AI platform offering real-time speech recognition, synthesis, and fully managed voice agent APIs.
110Deepgram is a leading voice AI platform that provides developers with high-accuracy, cost-effective real-time speech-to-text (STT), text-to-speech (TTS), and a unified voice agent API. Its Nova series models outperform competitors in both accuracy and speed; Aura-2 TTS offers latency under 200 milliseconds, and the voice agent API is priced at just $4.50 per hour. Supporting both cloud and self-hosted deployments, the platform is trusted by over 200,000 developers.
2026-07-11Listnr—A tool featuring over 1,000 hyper-realistic AI voices and voice cloning capabilities, reaching a global audience across 142 languages.
106Listnr is a powerful AI voice generation and cloning platform offering over 1,000 realistic AI voices and supporting more than 142 languages and accents. Key features include rapid 30-second voice cloning, an AI video generator, and podcast hosting services, enabling content creators, marketers, and businesses to easily produce multilingual audio and video content.
2026-07-11Ciniaoniao Voiceover—a permanently free AI voiceover tool; over 200 voices transform text into expressive speech in seconds.
106Ci Niao Voiceover is a free, AI-powered text-to-speech application featuring over 200 voice options and support for Mandarin, Cantonese, English, and various dialects. It offers nearly 300 distinct voices and more than ten emotional styles, catering to use cases such as short-video voiceovers, film and TV commentary, and audiobooks. The service is accessible across multiple platforms, including the web, mobile apps, and mini-programs.
2026-07-11Voice AI—a real-time voice-changing and voice-cloning tool; a "voice skin" for game streaming and content creation.
106Voice AI is a real-time voice changing and cloning platform for Windows that utilizes virtual microphone technology to apply voices in real-time across scenarios such as gaming, live streaming, and meetings. It offers hundreds of voice options and custom cloning capabilities, supports integration with applications like Discord, Zoom, and OBS, and serves millions of creators and gamers worldwide.
2026-07-11Voicemod—A real-time voice changer and soundboard that brings gaming voice chat and live streams to life.
106Voicemod is real-time voice-changing software featuring an extensive library of AI voices and sound effects, allowing users to instantly alter their voices or play humorous sound effects during gaming, live streaming, and voice calls. It integrates seamlessly via virtual microphone technology, enabling use without the need for additional hardware.
2026-07-10TTSMaker—A free online text-to-speech tool; generate over 200 voice styles with a single click.
103TTSMaker is a free online AI text-to-speech tool that supports multiple languages and over 200 voice styles. No registration is required; simply enter text to generate natural, realistic speech, adjust speed and volume, and download MP3 files. Ideal for video voiceovers, audiobook production, and educational content creation, it allows text to easily "come to life" with a voice.
2026-07-11Uberduck—an all-in-one voice studio where AI speaks, sings, and raps, with over 5,000 voices at your disposal.
103Uberduck is an AI-powered platform for voice and music creation, offering features such as text-to-speech, voice cloning, text-to-singing/rapping, and AI music generation. With a library of over 5,000 expressive voices and support for more than 70 languages and hundreds of musical styles, the platform enables creators, musicians, and marketers to rapidly produce professional-grade audio content.
2026-07-11Krisp—AI noise cancellation and meeting assistant for clearer, more efficient remote communication.
103Krisp is an AI-powered voice platform offering two-way noise cancellation, AI accent transformation, and an intelligent meeting assistant. It provides real-time transcription, generates summaries and action items, supports multilingual translation, and integrates with meeting tools such as Zoom and Teams. Free trials are available for both individual and team plans.
2026-07-11Murf AI—From voiceover studio to real-time voice agent: Generate professional-grade voices with AI.
102Murf AI is an AI voice generation and conversational platform offering over 200 ultra-realistic voices, support for more than 20 languages, and voice cloning capabilities. Its Falcon TTS API enables real-time voice agent deployment with an industry-leading 55ms latency, while the Studio editor allows for audio-video synchronization, making it suitable for applications such as video voiceovers, course creation, and intelligent customer service.
2026-07-11MiniMax Audio—an ultra-realistic large-scale speech model capable of everything from 10-second voice cloning to support for over 40 languages, enabling AI to speak with the warmth of a real person.
102MiniMax Audio is an AI voice platform under MiniMax that offers features such as text-to-speech, voice cloning, and music generation. Its proprietary Speech series models support over 40 languages, ultra-long text, emotional expression, and low-latency interaction, and have been adopted by leading global platforms and products such as LiveKit, Pipecat, Gaotu, and Ximalaya.
2026-07-11Wondercraft—an AI studio that drives video and audio creation through dialogue—cuts professional content production time from weeks to minutes.
101Wondercraft is an AI-powered platform for video and audio creation. Through its built-in AI agent, "Wonda," users can produce and edit content—such as podcasts, advertisements, and training videos—using natural language conversations. The platform integrates ElevenLabs' hyper-realistic voice technology, supports over 30 languages and voice cloning, and facilitates team collaboration; it has been adopted by organizations including Spotify, Amazon, and the World Bank.
2026-07-11Makefun AI - Free & Unrestricted AI Video & Image Generator with Text-to-Video, Face Swap, Voice Clone
100Makefun AI is a powerful free AI tool that supports text-to-image, image-to-video, face swap, head swap, cloth swap, voice cloning, avatar generation, and more. It features cutting-edge models like Seedream, Wan, Kling, Sora, etc. With a one-time payment, you get unlimited access to all premium features without subscription. Perfect for content creators, designers, and marketers to generate creative visual content quickly.
2026-07-27GPT-4o—OpenAI's all-around multimodal flagship model, featuring real-time voice and video interaction, and freely available to all users.
100GPT-4o (“o” stands for Omni) is OpenAI's next-generation flagship multimodal large model, released in May 2024. It enables real-time inference across text, audio, images, and video, with an audio response time as fast as 232 milliseconds, approaching human conversational reaction speed. It possesses groundbreaking capabilities such as emotion perception, real-time translation, and image generation. Its API is twice as fast as GPT-4 Turbo, costs only half the price, and is available to all free ChatGPT users.
2026-07-14Suno—an AI music powerhouse that generates full songs from a single sentence, a new creative favorite for over 100 million people worldwide.
99Suno is an AI music generation platform that creates complete songs—featuring vocals, lyrics, and instrumentation—within 30 seconds based on any input idea. Version 5.5, released in March 2026, introduced features such as voice cloning, custom models, and personalized learning. The platform supports the generation of extended tracks up to eight minutes long and covers a vast array of musical styles. Free users can generate 50 songs daily, while paid plans start as low as $8 per month. Suno boasts over 100 million global users, 2 million paying subscribers, and an annual recurring revenue of $300 million.
2026-07-11Designs.ai—an all-in-one creative studio where you assemble AI models like LEGO bricks; it covers everything from logos to videos and locks in your brand style with a single click.
99Designs.ai is an all-in-one creative platform integrating various AI models and offering a comprehensive suite of tools—including logo generation, video production, image creation, copywriting, and voice synthesis. Its core feature is "brand consistency": users can extract brand visual guidelines via a single URL, ensuring that all output automatically aligns with the brand's identity. Requiring no prior design experience, the platform is suitable for individual creators, marketing teams, and enterprises alike.
2026-07-10Moyin (Moyin Workshop)—an AI voiceover powerhouse featuring over 800 voices and 1,000 styles, trusted by creators of short videos and audiobooks.
97Moyin (Moyin Workshop) is an AI voice synthesis platform under Mobvoi. Powered by the proprietary "Sequence Monkey" (Xulie Houzi) large model and a fifth-generation TTS engine, it offers over 800 voice profiles, more than 1,000 styles, and nearly 20 fine-tuning features. The platform supports multiple languages and dialects, voice cloning, cloud-based video editing, and multi-user collaboration; it is widely used in applications such as short videos, audiobooks, and film/TV commentary, having served over 6 million users to date.
2026-07-11Beatoven.ai—an AI music generator built for creators; create emotion-driven soundtracks and ensure copyright never stands in the way of your creativity.
97Beatoven.ai is an AI-powered, royalty-free music generation platform specializing in creating emotive soundtracks for videos, podcasts, games, and more. It generates unique background music from text, images, or videos, offering customization across 16 moods and various styles. Its Maestro model is trained on fully licensed data and provides royalty sharing for musicians, having already helped creators worldwide generate millions of tracks.
2026-07-11LOVO AI — A lifelike TTS platform with an integrated AI video editor, featuring over 500 voices and more than 100 languages.
97LOVO AI is a high-fidelity text-to-speech and video creation platform. Its core product, Genny, integrates voice generation with online video editing, offering over 500 voices and 100 languages, alongside features such as AI scriptwriting and automatic subtitling. Supporting voice cloning and team collaboration, the platform is ideal for content creators, marketers, and educators.
2026-07-11Langlang Voiceover—a permanently free AI voiceover tool featuring over 1,100 voice talents, support for 80+ languages, and nearly 20 fine-tuning options.
96Langlang Voiceover is a permanently free AI voiceover platform featuring over 1,100 AI voices, support for more than 80 languages, and over 10 emotional styles. It offers nearly 20 fine-tuning features—such as continuous reading, pauses, handling of polyphones (characters with multiple pronunciations), localized speed adjustment, and multi-speaker narration—allowing for word-by-word customization of the voice output. The platform includes a built-in library of royalty-free (CC0) background music and provides productivity tools for subtitle generation, batch processing, and text extraction, making it ideal for short-video voiceovers, audiobook production, advertising, and more.
2026-07-11Lumina—An AI "second brain" for your browser; unlock ChatGPT, Claude, and Gemini using your own API keys.
95Lumina is an AI browser assistant that allows users to seamlessly access top-tier models—such as ChatGPT, Claude, and Gemini—directly from the sidebar using their own API keys. It supports intelligent analysis of the current webpage, Q&A based on screenshots, and voice input. With all data stored locally and a strong focus on privacy, it is an ideal tool for students, developers, and researchers.
2026-07-14Supertone Shift — Put the voices of professional voice actors into your throat: an ultra-low-latency, real-time AI voice changer from HYBE.
95Supertone Shift is a real-time AI voice-changing software developed by Supertone (a subsidiary of HYBE). It is renowned for its industry-leading low latency—as low as 47ms—and a lightweight design that operates without the need for a dedicated GPU. The software offers a vast library of AI voice personas created by professional voice actors and supports fine-tuned adjustments to parameters such as pitch, dynamic range, and reverb. Ideal for scenarios ranging from multiplayer gaming and VTuber streaming to video voiceovers, it is highly popular among creators worldwide, particularly in Japan.
2026-07-11AssemblyAI—the speech recognition API of choice for developers—build next-generation speech AI applications with industry-leading accuracy.
93AssemblyAI offers a production-grade speech-to-text API that supports 99 languages with over 93.3% word accuracy and pricing as low as $0.15 per hour. The platform integrates speech understanding capabilities—such as speaker diarization, PII identification, and summarization—and supports real-time streaming transcription with latency as low as 300ms; it is used by over 200,000 developers to build applications such as voice agents, meeting recorders, and conversation analytics tools.
2026-07-11There are 52 more tools not shown. Try more specific keywords.
AI News (7)
I use voice AI to turn text into songs, which provides an additional source of income (with
322Let me be honest with you first. I am someone who sings off key, has zero knowledge of music theory, and can't even understand sheet music. Whoever used to say 'you make money by making music', I must think there's something wrong with him. But that's the way I am, now I can earn more money every
Best AI Voiceover Tools Compared: Top 5 Value Picks in 2026
38Introduction: Is AI Voiceover Really Worth Dedicated Study? To be honest, when I first started exploring AI voiceover tutorials, I was a bit resistant. I assumed AI-generated voices would sound roboti...
AI Voiceover Tutorial & App Guide: 2026 Industry Playbook with 10 Success Case Studies
37Industry Background: When Voice No Longer Requires a "Human" to Speak To be honest, I've been involved in the AI voiceover space for over three years now. I still remember back in 2022, AI voiceovers ...
AI Voiceover Tutorials Reviewed: 2026 Performance, Cost & Use Cases Compared
35Comprehensive Review of AI Voiceover Tutorials: A Three-Dimensional Comparison of Performance, Cost, and Use Cases in 2026 — Essential Reading for Tool Selection Hey folks, fellow creators, short-vid...
2026 AI Voiceover Tutorial: 5 Best Cost-Effective Tools Reviewed – Avoid These Mistakes
29Preface: In This Day and Age, Not Using AI Voiceover Is a Huge Missed Opportunity Folks, don't scroll away just yet! Let me ask you a tough question: When creating short videos, running a Bilibili cha...
AI Voiceover Masterclass: From Beginner to Pro in 7 Days - Step-by-Step Guide to Mastering Core AI Skills
20Background: Why Everyone Should Learn AI Voiceover Now? Folks, let me be upfront: AI voiceover is no longer some exotic novelty, but those who can truly master it and make it shine are still few and ...
Best Multimodal AI Templates: 2026 Guide with Tips & Pitfalls to Avoid
11Opening Remarks: What Exactly Is Multimodal AI? Hey folks, don't scroll away just yet! I know you've been bombarded with all sorts of AI tools lately—text-to-image, image-to-video, voice cloning... it...
Career Skills (9)
Eldercare Exercise: Voice-Guided Bedridden Rehabilitation & Home Assistant Integration
A lightweight, TTS-guided exercise program for bedridden or seated elderly. Features 3 levels (super light, light, moderate), integrates with Home Assistant for cron scheduling, motion detection, SOS conflict prevention, multi-elder support, and daily reporting. Safe for 90+ with physician approval required.
Eldercare Exercise — Voice-Guided Bedridden Workout with Auto Timer and Level Adjustment
A lightweight exercise skill designed for elderly individuals who are bedridden or can sit. It uses TTS to guide each movement step by step, with automatic counting, rest breaks, and three difficulty levels (bedridden, sitting, standing with support). Supports daily cron at 9 AM or chat triggers, integrates with Home Assistant media player for voice output. Disabled by default; requires family activation and medical clearance.
Inspect Trash Bin Skill
This is a skill for inspecting the status of trash bins, built as one of the default skills for the Askme voice AI assistant. It helps users quickly understand the current state of a trash bin, such as whether it is full or needs cleaning. Suitable for smart home, office, or public space management scenarios, users can obtain trash bin information via voice commands, improving daily management efficiency. The skill is developed based on the askme project on GitHub and supports custom configuration and extension.
Multimodal Repair Skill
Multimodal Repair Skill is a core AI skill module in the SmartEstate intelligent real estate system, supporting home device fault diagnosis and repair guidance through multiple input methods such as text, images, and voice. This skill integrates multimodal AI models, capable of recognizing user-uploaded device photos, problem descriptions, or voice commands, automatically analyzing fault causes, and generating detailed repair steps or recommending professional services. It is suitable for smart home maintenance, property management, facility repair reporting, and other scenarios, significantly improving repair efficiency and user experience.
Cleaning System Skill
The Cleaning System Skill is a core skill of the Wayland AI agent for smart homes, focusing on the management and automation of household cleaning tasks. This skill perceives environmental conditions, reasons about cleaning needs, and executes corresponding cleaning operations such as sweeping, mopping, and vacuuming. It seamlessly integrates with smart home devices like robot vacuums and smart mops, automatically planning cleaning routes and adjusting strategies based on room layout and dirt levels. Users can initiate cleaning tasks via voice or text commands, and the skill provides real-time progress updates and maintains cleaning history. Suitable for indoor environments like homes and offices, it aims to improve cleaning efficiency, reduce manual intervention, and allow users to enjoy smarter, cleaner living spaces.
Roomba Control Skill
This skill enables you to control your Roomba robot vacuum directly via natural language commands. You can start, pause, stop cleaning, send it back to the charging dock, or check its current status. It's ideal for Roomba owners who want quick voice or text control without opening another app. Built on the openclaw platform, it integrates with Roomba's API for command execution and status feedback, making smart home control more convenient.
F-36 Plant Floor Operations: Integrated Quality Testing and Production Management
F-36 Plant Floor Operations is a skill designed for plant floor operators, integrating machine testing, disposition, shift management, vendor review, and compliance monitoring via CLI. It is optimized for noisy environments with gloved hands, supporting voice-driven workflows and Stream Deck button interactions.
Medical Scribe Dictation & Academic Writing Assistant: Fast and Accurate Medical Voice Transcription Tool
This skill is tailored for medical research and clinical settings, capable of transcribing doctors' dictations (case notes, surgical reports, diagnostic summaries) into structured text, and supporting academic writing with terminology proofreading, format optimization, and citation management. It integrates NLP and medical knowledge bases to significantly improve medical documentation efficiency and accuracy, suitable for hospitals, research institutions, and medical schools.
Nutrition Tracking Skill - Self-hosted Family Health Assistant for Diet Logging and Analysis
A nutrition tracking skill designed for the Wellness Advisor persona in StewardOS. It enables manual or voice-based meal logging, auto-calculates calories, macros, and micronutrients, tracks progress against personal goals, generates visual reports, and integrates with MCP servers and systemd runtime for a complete self-hosted health data pipeline.