Search results for "Multimodal"
Found 18 results (16 tools · 1 articles · 1 skills). Sorted by relevance by SeoAIu.
AI Tools (16)
OpenCompass: An Open Source Large Model Comprehensive Evaluation System and Sinan Open Platform
177OpenCompass is an open-source large model evaluation system launched by Shanghai Artificial Intelligence Laboratory, providing one-stop evaluation services for large language models, multimodal models, and scientific intelligence models. The platform supports one click distributed evaluation of
2026-05-31Alibaba's M6—the world's first 10 trillion-parameter multimodal large model—consumes only 1% of the energy of GPT-3.
110M6 is a general-purpose multimodal large model developed by Alibaba DAMO Academy. In October 2021, it became the world's first AI pre-trained model with 10 trillion parameters. Based on the self-developed Whale distributed framework, it was trained in just 10 days using only 512 V100 GPUs, with energy consumption only 1% of GPT-3. It supports tasks such as image and text generation, visual question answering, and poetry creation, and has been deployed in more than 40 scenarios, including Tmall virtual anchors and Taobao search. It is the core predecessor of the Tongyi Thousand Questions large model.
2026-07-02LLaMA—Meta's open-source AI model revolution, with performance comparable to GPT-4 and completely free for commercial use.
94LLaMA is an open-source family of large language models launched by Meta, evolving from LLaMA 1 in February 2023 to multimodal LLaMA 4. The 70B parameter version boasts performance close to GPT-4, while the 405B version rivals GPT-4o. It is completely open-source and commercially viable. Enterprises can download and deploy it on their own servers without data sharing, meeting compliance requirements in industries such as finance and healthcare. With over 1.2 billion downloads globally, it is suitable for developers and enterprises requiring private deployment and deep customization.
2026-07-02360AI Search – an AI search engine that can "think slowly" for you, so you no longer need to flip through dozens of pages to find information.
81360AI Search (Nano AI Search) is an AI-native search engine launched by 360, supporting multimodal input including text, voice, photos, and videos. It features a built-in "slow thinking mode," capable of calling up three or more large models at once and generating a 5,000-word in-depth report after reading 250,000 documents. It supports extracting mind maps from 100-page PDF/Word documents in one minute, extracting key points from one hour of video in one minute, and providing a one-stop AI tool for webpage AI summarization, PPT generation, and academic source tracing. In July 2024, user visits exceeded 90 million, transforming search from "providing links" to "providing direct answers."
2026-07-02360 Smart Drawing – an AI drawing tool equipped with a massive model containing hundreds of billions of parameters, capable of generating thousands of styles with a single click.
74360 Smart Drawing is an AI image creation platform launched by 360 based on its self-developed "360 Brain" large-scale model with hundreds of billions of parameters. Integrating the 360CV large-scale model and multimodal technology, it provides full-scene AI drawing and retouching functions, including text-to-image, image-to-image, doodle-to-image, partial redrawing, AI portrait, and LoRA model training. It has built-in models for thousands of styles, including traditional Chinese style, realistic, animation, and 3D, and supports bilingual (Chinese and English) prompts. As a member of the 360 AI ecosystem, it has been integrated into the entire 360 platform ecosystem.
2026-07-04ModelScope – China's largest and most active open-source AI model platform, boasting a "model-as-a-service" ecosystem with over 30 million developers.
64ModelScope is a leading open-source AI model platform in China, initiated by Alibaba DAMO Academy in conjunction with the CCF Open Source Development Committee. Adhering to the "Model as a Service" (MaaS) concept, the community has gathered over 200,000 open-source models, 30 million developers, and over 500 contributing organizations. It provides end-to-end services from model experience, download, fine-tuning, training to deployment, covering LLM, multimodal, speech, and AIGC fields. As a core infrastructure for AI developers in China, it is promoting the democratization of AI technology through "open source."
2026-07-14GPT-4o—OpenAI's all-around multimodal flagship model, featuring real-time voice and video interaction, and freely available to all users.
59GPT-4o (“o” stands for Omni) is OpenAI's next-generation flagship multimodal large model, released in May 2024. It enables real-time inference across text, audio, images, and video, with an audio response time as fast as 232 milliseconds, approaching human conversational reaction speed. It possesses groundbreaking capabilities such as emotion perception, real-time translation, and image generation. Its API is twice as fast as GPT-4 Turbo, costs only half the price, and is available to all free ChatGPT users.
2026-07-14Google DeepMind Gemma – A lightweight, open-source, commercially viable next-generation AI model family
56Gemma is an open-source AI model family built by Google DeepMind based on Gemini technology. It offers various parameter scales from 2B to 31B, supports multimodal understanding including text, images, and audio, and supports over 140 languages. It is licensed under the Apache 2.0 license and can be used commercially for free. Known for its lightweight and high performance, Gemma can run locally on consumer hardware and has been downloaded over 400 million times, making it a mainstream choice for developers building AI applications.
2026-07-14ThinkAny—an AI search engine based on RAG, a new way to obtain information born over a weekend.
56ThinkAny is an AI search engine based on RAG (Retrieval Augmentation) technology. By aggregating high-quality online content and combining it with AI-powered question answering, it provides direct and accurate answers rather than lists of links. It supports multimodal search, mind-map-style summaries, and has spawned the desktop AI agent WorkAny. Developed by its creators in a single weekend, it gained 170,000 users within three months of its launch.
2026-07-14Yuanjing - AI-Powered Rapid Content Creation Engine for Script Generation and Multimodal Storyboarding
47Yuanjing is an AI-powered rapid content creation engine developed by Zelins Intelligence, specializing in efficient and professional content production. It integrates creative script generation, multimodal storyboard design, and one-click video creation, catering to various scenarios such as short videos, advertisements, and promotional films. Users can quickly transform ideas into high-quality video content, significantly enhancing creative efficiency. Whether for individual creators or professional teams, Yuanjing provides powerful AI assistance to make video production simpler and smarter.
2026-07-17360 Brain - Multimodal AI Model, Safe and Trustworthy AI Assistant
44360 Brain is a multimodal AI model developed by 360, integrating language, image, and video capabilities. It offers intelligent dialogue, content generation, and knowledge Q&A. With a focus on 'people-centric, safe and trustworthy', it excels in data privacy and content security, suitable for personal and enterprise users in creative writing, information retrieval, and data analysis. Supports multi-turn conversation and real-time interaction.
2026-07-22Gemini 3.5 Multimodal AI Model for Efficient Coding, Knowledge Work, and Multimodal Tasks
43Gemini 3.5 is a next-generation multimodal AI model from Google DeepMind, optimized for efficient coding, knowledge reasoning, and multimodal task processing. It reduces computational costs while maintaining high performance, making it ideal for fast iteration and resource-sensitive applications.
2026-07-22Qwen AI Assistant - Versatile Chatbot for Content Creation & Problem Solving
38Qwen Studio is a free AI chat platform by Alibaba Cloud, powered by the Qwen large language model. It offers intelligent Q&A, copywriting, code generation, data analysis, and creative planning. With a clean interface and fast responses, it supports multi-turn conversations and multimodal understanding, ideal for study, work, and daily life.
2026-07-22Vozo AI Video Translator, Dubbing & Lip Sync
38Vozo AI is an AI-powered video localization tool that supports translation, dubbing, lip sync, subtitles, and on-screen text in over 160 languages. Built for creators, marketers, and educators, it leverages multimodal AI to understand scenes, context, and tone, delivering natural, fluent, and human-level accurate video localization, significantly improving efficiency and reducing costs.
2026-07-17Tiangong AI Office Agent, Deep Research, One-Click Doc/PPT/Sheet Generation
37Tiangong is a super agent with powerful DeepResearch capabilities, integrating advanced multimodal understanding and deep retrieval analysis. It enables one-click generation of AI documents, AI PPTs, and AI spreadsheets, efficiently handling office and learning tasks. It also supports creative content creation in forms such as web HTML, images, videos, audiobooks, and picture books, inspiring unlimited creativity. Whether you are an office worker, researcher, student, or KOL, Tiangong delivers research-grade, professional, and consulting-level results, helping you focus on thinking and unleash creativity.
2026-07-22Saylo AI - Immersive Role-Play Chat and Interactive Story Creation
36Saylo AI is an immersive platform for virtual character chat and interactive story creation. Users can engage in natural conversations with various preset or custom AI characters and participate in AI-driven narrative adventures. With multimodal interaction support, rich character settings, and branching storylines, it empowers everyone to become a storyteller and enjoy a unique virtual social and entertainment experience.
2026-07-22