AI open source project

Total 1020 articles posts
LiberSonora:有声书字幕提取与多语言翻译,有声小说转录为多语言

LiberSonora: Audiobook Subtitle Extraction and Multilingual Translation, Audiobook Transcription into Multiple Languages

General Introduction LiberSonora, which means "free sound", is a powerful AI-enabled open source audiobook toolset. The toolset supports intelligent subtitle extraction, AI title generation, multi-language translation, etc., and is capable of batch offline processing under GPU acceleration.LiberSo...
2yrs ago
090.6K
CoT-Lab:探索人机协作迭代思考的实验性对话工具

CoT-Lab: an experimental dialog tool for exploring iterative thinking about human-computer collaboration

CoT-Lab is an experimental interface for exploring a new paradigm of human-computer collaboration. Based on Cognitive Load Theory and Active Learning Principles, CoT-Lab facilitates deep cognitive alignment between humans and Artificial Intelligence (AI) through the creation of "thinking partner" relationships. The program aims to...
2yrs ago
090.4K
MOFA Video:运动场适配技术将静态图像转换为视频

MOFA Video: Motion Field Adaptation Technology Converts Still Images to Video

General Introduction MOFA-Video is a state-of-the-art image animation generation tool that utilizes generative motion field adaptation techniques to convert static images into dynamic videos. The project was developed in collaboration with the University of Tokyo and Tencent AI Lab, and will be presented at the 2024 European Conference on Computer Vision (E...
2yrs ago
090.4K
MM-EUREKA:探索视觉推理的多模态强化学习工具

MM-EUREKA: A Multimodal Reinforcement Learning Tool for Exploring Visual Reasoning

Comprehensive Introduction MM-EUREKA is an open source project developed by Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University and other parties. It extends textual reasoning capabilities to multimodal scenarios through rule-based reinforcement learning techniques to help models process image and textual information. The core of this tool...
2yrs ago
090.3K
zChunk:基于Llama-70B的通用语义分块策略

zChunk: a generic semantic chunking strategy based on Llama-70B

Comprehensive Introduction zChunk is a novel chunking strategy developed by ZeroEntropy that aims to provide a solution for generic semantic chunking. The strategy is based on the Llama-70B model, which optimizes the chunking process of documents by prompting for chunks to be generated, ensuring that information retrieval is maintained at a high...
2yrs ago
090.2K
Marco-o1:基于Qwen2-7B-Instruct微调的开源版OpenAI o1模型,探索开放式推理模型,解决复杂问题

Marco-o1: An Open Source Version of the OpenAI o1 Model Based on Qwen2-7B-Instruct Fine-Tuning to Explore Open Inference Models for Solving Complex Problems

Comprehensive Introduction Marco-o1 is an open reasoning model developed by Alibaba International Digital Commerce Group (AIDC-AI) to solve complex real-world problems. The model combines Chain of Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), and innovative reasoning strategies...
2yrs ago
089.9K
DisPose:生成人体姿态精准控制的视频,创作跳舞的小姐姐

DisPose: generating videos with precise control of human posture, creating dancing ladies

General Introduction DisPose is an innovative open source artificial intelligence project focused on controlled character image animation generation. Developed by a team of researchers and open-sourced on GitHub, the project uses advanced deep learning techniques to achieve precise character animation control by decomposing skeletal pose information.D...
2yrs ago
089.7K
LettuceDetect:检测RAG系统幻觉的高效工具

LettuceDetect: an efficient tool for detecting hallucinations in the RAG system

Comprehensive Introduction LettuceDetect is a lightweight open-source tool developed by KRLabsOrg that specializes in detecting hallucinatory content generated in Retrieval Augmented Generation (RAG) systems. It identifies responses that are not supported by the context by comparing the context, the question, and the answer...
2yrs ago
089.4K
Pyramid Flow:快手推出的开源版

Pyramid Flow: an open source version of "Kringle" launched by Racer, based on SD3 and running on GPUs of less than 8GB (one-click deployment version)

Comprehensive Introduction Pyramid Flow is an efficient autoregressive video generation method based on the Flow Matching technique. The method achieves higher computational efficiency in generating and decompressing video content by interpolating between different resolutions and noise levels...
2yrs ago
089.3K
muAgent:由 LLM 和 EKG(行业知识)驱动的全新Agent编排框架

muAgent: A New Agent Orchestration Framework Driven by LLM and EKG (Industry Knowledge)

General Introduction muAgent is an innovative multi-intelligentsia framework developed by Ant Group. The framework collaborates with multi-intelligentsia, function calls, code interpreters and other technologies through canvas drag-and-drop and simple text writing to help users execute various complex standard operating procedures (SOPs) under human guidance...
2yrs ago
088.4K
Kheish:多角色智能体,审查、验证和格式化输出以生成高质量结果

Kheish: multi-actor intelligences that review, validate and format output to produce high quality results

Comprehensive Introduction Kheish is an open source multi-role agent designed for Large Language Model (LLM) tasks that require structured, step-by-step collaboration.Kheish is more than just a simple coordinator, it is an intelligent agent in its own right, requesting modules on demand, integrating user-reversal...
2yrs ago
088.1K
中文基于满血 DeepSeek-R1 蒸馏数据集,支持中文R1蒸馏SFT数据集

Chinese based full-blooded DeepSeek-R1 distillation dataset, supports Chinese R1 distillation SFT dataset

Comprehensive Introduction The Chinese DeepSeek-R1 distillation dataset is an open source Chinese dataset containing 110K pieces of data designed to support machine learning and natural language processing research. The dataset is released by Cong Liu's NLP team. The dataset contains not only mathematical data, but also a large number of general types...
2yrs ago
088K
SciToolAgent:整合500+科研工具,自动化研究科研任务的智能体

SciToolAgent: Integration of 500+ research tools and automation of research and scientific tasks for intelligent bodies

Comprehensive Introduction SciToolAgent is an open source tool platform developed by the Innovation Center of Zhejiang University in Hangzhou (HICAI-ZJU). It integrates more than 500 scientific tools through knowledge graph (SciToolKG) and big language modeling technologies to help researchers deal with...
2yrs ago
088K
HN中文播客:自动抓取热门科技文章,AI生成中文总结并转换为播客

HN Chinese Podcast: Automatically grab popular tech articles, AI-generated Chinese summaries and convert them to podcasts

General Introduction The Hacker News Chinese Podcast project is an innovative platform based on AI technology, aiming to automatically grab popular articles on Hacker News every day and generate Chinese summaries and podcast content through AI. The project is led by ccbikai ...
2yrs ago
087K
AgentClientDemo:演示智能体运行过程的Python客户端,提供直观的图形用户界面

AgentClientDemo: a Python client that demonstrates the process of running an intelligent body, providing an intuitive graphical user interface

Comprehensive Introduction AgentClientDemo is a comprehensive Python project that integrates intelligent (Agent) and client (Client) functionality. The project is based on the PyQt framework and provides an intuitive and easy-to-use graphical user interface (G...
2yrs ago
086.8K
FoloUp:开源AI语音面试平台,生成定制面试题并进行智能分析

FoloUp: Open Source AI Voice Interview Platform Generates Customized Interview Questions and Performs Intelligent Analysis

General Introduction FoloUp is an open source platform that specializes in AI-powered voice interview solutions for enterprises. With FoloUp, enterprises can quickly generate customized interview questions for job descriptions and conduct natural conversational interviews with AI. The platform also provides detailed interview analysis...
2yrs ago
086.7K
Vision is All You Need:使用视觉语言模型构建智能文档检索系统(Vision RAG)

Vision is All You Need: Building an Intelligent Document Retrieval System Using Visual Language Models (Vision RAG)

Comprehensive Introduction Vision-is-all-you-need is an innovative visual RAG (Retrieval Augmented Generation) system demonstration project that breaks new ground in applying Visual Language Modeling (VLM) to the document processing domain. Unlike traditional text chunking methods, the system directly makes...
2yrs ago
086.7K
混元Turbo S:腾讯推出的快思考大模型(开放申请)

Hybrid Turbo S: Tencent's Big Model of Fast Thinking (open for applications)

Comprehensive Introduction Tencent Hybrid Turbo S is a new generation of Tencent's self-developed fast-thinking model, which has been launched on Tencent Cloud's official website, and will be officially released on February 27th, 2025. It is different from traditional slow thinking models (such as Deepseek R1 and Mixed Meta T1), and can realize "second reply", spit...
2yrs ago
086.7K
OmniParse:从文档/多媒体中提取任何非结构化数据解析为结构化数据

OmniParse: extract any unstructured data from documents/multimedia and parse it into structured data

Comprehensive Introduction OmniParse is a powerful data parsing and optimization platform designed to convert any unstructured data into structured, actionable data optimized for GenAI (Generative Artificial Intelligence) framework. Whether you are working with documents, tables, images, videos, audio files or...
2yrs ago
086.3K
Swarm:学习轻量级多智能体系统的实验性教学项目(OpenAI示例)

Swarm: an experimental pedagogical program for learning lightweight multi-intelligent body systems (OpenAI example)

General Introduction Swarm is an experimental educational framework developed by OpenAI to explore lightweight, controlled, and easy-to-test interfaces for multi-agent systems. The framework is primarily used to demonstrate handoffs and routine patterns between agents to help developers understand and implement the coordination and execution of multi-agent systems...
2yrs ago
086.1K
wdoc:从海量、多源文档中检索内容并总结知识

wdoc: retrieve content and summarize knowledge from massive, multi-source documents

Comprehensive Introduction wdoc is a powerful RAG (Retrieval Augmentation Generation) system designed for processing and analyzing large and diverse documents. It is capable of retrieving from a wide range of document types, including PDFs, web pages, YouTube videos, audio files, etc. wdoc is particularly well suited for processing...
2yrs ago
086K
Claude生成深度研究报告的MCP服务

Claude's MCP service for generating in-depth research reports

Comprehensive Introduction MCP Server Deep Research is an open source tool that automatically generates structured research reports for complex problems through artificial intelligence and web search. Users enter a research question, and the tool breaks down the question, searches for authoritative information, assesses source credibility...
1yrs ago
086K
Inbox Zero:轻松实现收件箱零邮件,借助 AI 帮助你对邮件进行归类、过滤、处理。

Inbox Zero: Easily achieve zero emails in your inbox, with the help of AI to help you categorize, filter, and process your emails.

General Description Inbox Zero is an open source email management app designed to help users quickly achieve inbox zero emails with an AI assistant. The app offers a variety of features including auto-replying, archiving, labeling and forwarding emails, managing and unsubscribing from newsletters, blocking cold emails, following...
2yrs ago
085.6K