Search results for "LLM"

Found 29 results (26 tools · 1 articles · 2 skills). Sorted by relevance by SeoAIu.

AI Tools (26)

FlagEval: The internationally authoritative large model evaluation system and Libra open platform

206
FlagEval: The internationally authoritative large model evaluation system and Libra open platform
AI Tools

FlagEval (Libra) is a large-scale model evaluation system and open platform initiated by Beijing Zhiyuan Artificial Intelligence Research Institute, aiming to establish scientific, fair, and open evaluation benchmarks and methods. The platform has innovatively constructed a three-dimensional

2026-05-31

AGI Eval: A large model evaluation community and authoritative third-party evaluation platform

161
AGI Eval: A large model evaluation community and authoritative third-party evaluation platform
AI Tools

AGI Eval is a large model evaluation community jointly created by top universities and institutions such as Shanghai Jiao Tong University, Tongji University, East China Normal University, and DataWhale, with the mission of "assisting evaluation and making AI a better partner for humanity". The

2026-05-31

OpenCompass: An Open Source Large Model Comprehensive Evaluation System and Sinan Open Platform

155
OpenCompass: An Open Source Large Model Comprehensive Evaluation System and Sinan Open Platform
AI Tools

OpenCompass is an open-source large model evaluation system launched by Shanghai Artificial Intelligence Laboratory, providing one-stop evaluation services for large language models, multimodal models, and scientific intelligence models. The platform supports one click distributed evaluation of

2026-05-31

CMMLU: Authoritative Chinese Language Model Knowledge Understanding Ability Evaluation Benchmark

148
CMMLU: Authoritative Chinese Language Model Knowledge Understanding Ability Evaluation Benchmark
AI Tools

CMMLU (Chinese Massive Multitask Language Understanding) is a large-scale language understanding benchmark designed specifically for the Chinese language context, covering 67 subject topics from beginner to advanced professional levels, including natural sciences, social sciences, engineering

2026-05-30

MMLU: The International Authoritative Benchmark for Multi Task Language Understanding Ability of

126
MMLU: The International Authoritative Benchmark for Multi Task Language Understanding Ability of
AI Tools

MMLU (Massive Multitask Language Understanding) is a large-scale multi task language understanding evaluation dataset jointly released by the University of California, Berkeley and other institutions. It covers 57 disciplinary fields, including humanities, social sciences, natural sciences,

2026-05-31

HELM: Stanford University led comprehensive evaluation framework for large language models and high

121
HELM: Stanford University led comprehensive evaluation framework for large language models and high
AI Tools

HELM (Holistic Evaluation of Language Models) is a comprehensive language model evaluation framework initiated by the Stanford University Center for Fundamental Model Research (CRFM), aimed at systematically evaluating large language models through multidimensional, standardized, and reproducible

2026-05-30

Wenxin Big Model: Baidu's Industry level Knowledge Enhancement Model that has been honed over the

112
Wenxin Big Model: Baidu's Industry level Knowledge Enhancement Model that has been honed over the
AI Tools

ERNIE is an industry level knowledge enhancement model independently developed by Baidu. Starting from the release of version 1.0 in 2019, it has undergone multiple generations of technological iterations and has established a three-level system of basic model, task model, and industry model. The

2026-06-03

SuperCLUE—A benchmark for Chinese large models: from foundational capabilities to AI agents, a single test reveals who is "swimming naked."

96
SuperCLUE—A benchmark for Chinese large models: from foundational capabilities to AI agents, a single test reveals who is "swimming naked."
AI Tools

SuperCLUE is a comprehensive evaluation benchmark for general-purpose Chinese large models, released by the CLUE team. It provides authoritative, multi-dimensional capability assessments and rankings for Chinese large models by regularly publishing monthly and semi-annual reports based on three key benchmarks—open-domain multi-turn dialogue, closed-domain objective questions, and anonymous head-to-head battles—as well as dimensions such as mathematical reasoning, code generation, and AI agents.

2026-07-10

StableLM—an open-source AI model from the same family as Stable Diffusion, capable of running with only 3B parameters and commercially viable.

96
StableLM—an open-source AI model from the same family as Stable Diffusion, capable of running with only 3B parameters and commercially viable.
AI Tools

StableLM is an open-source large language model series launched by Stability AI. It is trained on The Pile extended dataset, which contains 1.5 trillion tokens. It offers multiple parameter versions, including 1.6B, 3B, 7B, and 12B, and supports text generation and code writing. It is licensed under the CC BY-SA 4.0 open-source license, allowing free commercial use. Stable LM 2 12B outperforms Llama 2 70B in some benchmark tests. It is suitable for developers, researchers, and small and medium-sized enterprises for private deployment.

2026-07-02

Sapling AI Content Detector—a "truth serum" from a former Google researcher—can it see through the tricks of GPT-5 and Claude 4.5?

93
Sapling AI Content Detector—a "truth serum" from a former Google researcher—can it see through the tricks of GPT-5 and Claude 4.5?
AI Tools

Developed by a former Google researcher, Sapling AI Content Detector is a tool for identifying whether text is generated by AI. It claims a detection rate of over 97% for AI-generated content and a false positive rate of less than 3% for human text. It supports the latest models such as GPT-5, Claude 4.5, and Gemini 2.5, can process PDF and DOCX files, and offers a browser extension for convenient on-the-go detection on web pages.

2026-07-09

PubMedQA—an "AI benchmark" specifically designed for biomedical question answering; only by comprehending research papers can an AI truly pass the "Medical Turing Test."

84
PubMedQA—an "AI benchmark" specifically designed for biomedical question answering; only by comprehending research papers can an AI truly pass the "Medical Turing Test."
AI Tools

PubMedQA is the first question-answering dataset requiring reasoning over biomedical research texts; it was released in 2019 by institutions including the University of Pittsburgh. The task involves answering "Yes," "No," or "Maybe" questions based on PubMed abstracts. Comprising 1,000 expert-annotated samples and 211,000 artificially generated ones, the dataset aims to evaluate the ability of AI models to comprehend and reason about complex medical literature. It is widely used to benchmark the performance of large language models in the medical domain.

2026-07-10

AnythingLLM – Full-stack AI applications and private knowledge bases, enabling you to build your own ChatGPT with any model.

70
AnythingLLM – Full-stack AI applications and private knowledge bases, enabling you to build your own ChatGPT with any model.
AI Tools

AnythingLLM is a full-stack AI application that supports any commercial or open-source LLM. It features built-in RAG, AI agents, and no-code agent builders. It can run locally or be remotely hosted, intelligently converse with documentation, and requires no cumbersome setup. Supporting the MCP protocol, it offers Desktop and Docker deployment options, making it suitable for individuals and businesses to build private, fully functional AI assistants.

2026-07-14

MiTa AI Search – a free AI search engine that thinks and provides references, so you no longer need to flip through hundreds of pages of materials when doing research.

63
MiTa AI Search – a free AI search engine that thinks and provides references, so you no longer need to flip through hundreds of pages of materials when doing research.
AI Tools

MetaAI Search is an intelligent search engine launched by Shanghai MetaNet Technology Co., Ltd. based on its self-developed MetaLLM large-scale model. It offers three search modes: concise, in-depth, and research-oriented. It supports searches across the entire web, document libraries, academic resources, podcasts, and more. Search results automatically generate outlines and mind maps, with each result accompanied by references. It incorporates the DeepSeek R1 deep thinking model, supporting a "think first, then search" research mode. It is compatible with web pages, apps, and mini-programs, with a free daily search limit of 100 times. Suitable for students, researchers, and professionals, it transforms information retrieval from "providing links" to "providing direct answers."

2026-07-02

ModelScope – China's largest and most active open-source AI model platform, boasting a "model-as-a-service" ecosystem with over 30 million developers.

45
ModelScope – China's largest and most active open-source AI model platform, boasting a "model-as-a-service" ecosystem with over 30 million developers.
AI Tools

ModelScope is a leading open-source AI model platform in China, initiated by Alibaba DAMO Academy in conjunction with the CCF Open Source Development Committee. Adhering to the "Model as a Service" (MaaS) concept, the community has gathered over 200,000 open-source models, 30 million developers, and over 500 contributing organizations. It provides end-to-end services from model experience, download, fine-tuning, training to deployment, covering LLM, multimodal, speech, and AIGC fields. As a core infrastructure for AI developers in China, it is promoting the democratization of AI technology through "open source."

2026-07-14

Jan - Open Source Offline AI Assistant Tool for Running LLMs Locally with Privacy Protection

35
Jan - Open Source Offline AI Assistant Tool for Running LLMs Locally with Privacy Protection
AI Tools

Jan is a free and open-source AI assistant tool designed to run large language models locally on user devices without requiring internet connectivity. It supports multiple mainstream models (e.g., Llama, Mistral) and offers a clean, intuitive interface for easy model loading and management. Jan prioritizes privacy by processing all data entirely on-device, making it ideal for users concerned about data security. Suitable for personal learning, writing assistance, and lightweight development tasks, Jan delivers efficient and private AI experiences.

2026-07-15

Cohere - Enterprise AI Platform for Building RAG and Intelligent Search Applications

35
Cohere - Enterprise AI Platform for Building RAG and Intelligent Search Applications
AI Tools

Cohere is an enterprise-focused AI large language model platform that provides powerful natural language processing (NLP) capabilities, including text generation, semantic search, classification, and retrieval-augmented generation (RAG). Its key advantages include support for private deployment, data security and control, and optimized model accuracy and efficiency for enterprise use cases. It is widely used in intelligent customer service, knowledge base retrieval, document analysis, content generation, and more, helping enterprises quickly build LLM-based intelligent applications.

2026-07-15

Google PaLM 2—A next-generation AI language model with multilingual, strong reasoning, and coding capabilities.

31
Google PaLM 2—A next-generation AI language model with multilingual, strong reasoning, and coding capabilities.
AI Tools

PaLM 2 is Google's next-generation large language model, introduced at Google I/O 2023. It boasts enhanced multilingual capabilities (covering over 100 languages), logical reasoning and mathematical abilities, and code generation capabilities, offering four versions ranging from mobile-friendly Gecko to high-performance Unicorn. PaLM 2 has been integrated into over 25 Google products, including Bard, Google Workspace, Med-PaLM 2, and Sec-PaLM, serving as a core model driving Google's AI strategy.

2026-07-14

GitHub Copilot - AI Pair Programmer for Code Completion and Agent Workflows

27
GitHub Copilot - AI Pair Programmer for Code Completion and Agent Workflows
AI Tools

GitHub Copilot is an AI pair programmer by GitHub that provides intelligent code completion, explanations, edit suggestions, and autonomous agent execution across VS Code, JetBrains, CLI, and more. It integrates leading LLMs (OpenAI Codex, Claude) and supports custom agents. The GitHub Copilot app enables unified multi-agent workflow management. Enterprise plans allow customized knowledge bases, behavior policies, and service connections, boosting developer productivity by 94% for companies like Grupo Boticário.

2026-07-16

Beiji Jiuzhang Enterprise AI Data Insight Engine

23
Beiji Jiuzhang Enterprise AI Data Insight Engine
AI Tools

Beiji Jiuzhang is an enterprise-grade AI data insight engine powered by large language models and AI Agent technology, enabling natural language conversational data analysis. Users can ask questions via text or voice to generate data results with text and visuals, achieving 'conversation as analysis.' It features deep data analysis, AI-powered interpretation, multi-platform access, and trustworthy, controllable, and evolvable AI analysis. The product supports complex computations like year-over-year, correlation, attribution, and forecasting, serving dozens of leading enterprises in automotive, manufacturing, finance, retail, and FMCG industries, empowering business users to self-serve data analysis and generate smart reports for data-driven decisions.

2026-07-20

FARUI Legal AI Model by Alibaba Cloud - Legal Q&A, Document Generation & Case Analysis

20
FARUI Legal AI Model by Alibaba Cloud - Legal Q&A, Document Generation & Case Analysis
AI Tools

FARUI is a specialized large language model for the legal domain, developed by Alibaba Cloud. Designed for legal professionals, businesses, and the public, it offers capabilities for legal Q&A, reasoning legal applicability, recommending relevant precedents, assisting case analysis, generating legal documents, and retrieving legal knowledge. By leveraging AI, FARUI enhances efficiency and accuracy in legal work, enabling smart handling of contract review, case research, legal consulting, and more.

2026-07-20

WPS AI - AI-Powered PPT Generation, Document Writing & Spreadsheet Processing

19
WPS AI - AI-Powered PPT Generation, Document Writing & Spreadsheet Processing
AI Tools

WPS AI is an AI-powered application developed by Kingsoft Office, integrating large language models to enhance productivity. It enables one-click PPT outline generation, smart writing of weekly reports, annual review reports, social media copy, and more. It also offers rewriting, continuation, translation, and polishing features for documents. With spreadsheet data processing and analysis capabilities, WPS AI streamlines complex office tasks within the WPS Office ecosystem, delivering a fully intelligent workflow from creation to optimization.

2026-07-20

FormX.ai - AI-Powered Document Data Extraction Automation

19
FormX.ai - AI-Powered Document Data Extraction Automation
AI Tools

FormX.ai is an AI-powered document data extraction tool that automatically extracts structured data from invoices, receipts, bank statements, contracts, applications, and more. In just three steps—create an extractor, upload samples, and connect the API—you can seamlessly integrate it into your existing workflows. It supports switching between vision and LLM models, continuously improves accuracy with production data, and provides production-ready data with guardrails to reduce manual entry errors and boost business efficiency.

2026-07-20

DeepSeek - AI Chat, API Platform & Open-Source Large Language Models for AGI Research

17
DeepSeek - AI Chat, API Platform & Open-Source Large Language Models for AGI Research
AI Tools

DeepSeek, founded in 2023, is dedicated to advancing foundational AGI models and technologies. The platform offers free AI chat and API access, with open-source large language models including DeepSeek-LLM, DeepSeek-Coder, and DeepSeek-MoE. The latest DeepSeek-V4 preview delivers world-class reasoning performance and enhanced Agent capabilities, available on web, app, and API for seamless integration.

2026-07-22

Qwen-Character Star Dust - Role-Playing AI Dialogue Agent by Alibaba Cloud

16
Qwen-Character Star Dust - Role-Playing AI Dialogue Agent by Alibaba Cloud
AI Tools

Qwen-Character Star Dust is a role-playing AI dialogue agent developed by Alibaba Cloud, built on the Tongyi large language model. It enables rapid creation of unique personas and styles, widely used in role-playing, intelligent NPCs, virtual idols, and emotional companionship. Users can customize character traits, language styles, and backstories for natural multi-turn conversations. Ideal for game development, social entertainment, education, and more, it offers flexible character customization and API integration for immersive AI interactions.

2026-07-22

ZenMux - AI Multi-Model Hybrid Inference Platform

13
ZenMux - AI Multi-Model Hybrid Inference Platform
AI Tools

ZenMux is an innovative AI multi-model hybrid inference platform designed to optimize AI inference efficiency through intelligent routing and model composition. It supports simultaneous invocation of multiple large language models (LLMs), automatically selecting the optimal model or combination to balance cost, speed, and accuracy. Suitable for developers, AI researchers, and enterprises, it helps reduce API call costs, improve response speed, and enable more complex reasoning tasks. ZenMux provides flexible API interfaces and a visual dashboard for real-time monitoring and model switching.

2026-07-23

Unsloth AI Fine-Tuning Tool: Fast and Memory-Efficient LLM Training

8
Unsloth AI Fine-Tuning Tool: Fast and Memory-Efficient LLM Training
AI Tools

Unsloth is an open-source tool designed to accelerate fine-tuning of large language models, supporting Llama, Mistral, Gemma, and more. With double quantization and optimized kernels, it achieves up to 2x speed improvement and 80% memory reduction on consumer GPUs while maintaining model accuracy. No need for high-end hardware; free to use on Colab, ideal for individual developers and SMEs to quickly customize AI assistants.

2026-07-26

AI News (1)

Career Skills (2)