Latest AI Resources

Total 3143 articles posts
InternVL3.5 - 上海AI实验室开源的多模态大模型

InternVL3.5 - Shanghai AI Lab Open Source Multimodal Large Models

InternVL3.5 (Shusheng-Wanxiang 3.5) is an open source multimodal large model of the Shanghai Artificial Intelligence Laboratory, the model is fully upgraded in terms of general ability, reasoning ability and deployment efficiency, providing nine sizes of versions from 1 billion to 241 billion parameters, covering different resource demand scenarios, including thick...
11mos ago
066.9K
浙江大学免费PDF资料《大模型基础》 - 附下载链接

Free PDF of Fundamentals of Large Models from Zhejiang University - with download link

Fundamentals of Large Models provides an in-depth analysis of the core technologies and practical paths of Large Language Models (LLMs). Starting from the fundamental theory of language modeling, it systematically explains the principles of model design based on statistics, recurrent neural networks (RNN), and Transformer architecture, focusing on the three major big language model...
11mos ago
066.7K
Tizzy.ai - 百度推出的AI搜索应用

Tizzy.ai - AI search app launched by Baidu

Tizzy.ai is an AI intelligent search application launched by Baidu.Tizzy.ai is based on Baidu's big model technology, with powerful intelligent search functions, can quickly answer questions, deep thinking and assist in decision-making.Tizzy.ai has a simple interface, no ads and pop-ups, and the bottom of the guide...
1yrs ago
066.7K
Gemini CLI - 谷歌开源的编程Agent

Gemini CLI - Google Open Source Programming Agent

Gemini CLI is Google's open source AI programming tool based on incorporating the Gemini Big Model into the developer's endpoint to provide developers with powerful AI capabilities. The tool understands code, manipulates files, executes commands, and dynamically troubleshoots problems to help developers efficiently write generation...
1yrs ago
066.5K
绘想 - 百度推出的AI视频生成平台

Painting Thinking - AI Video Generation Platform Launched by Baidu

Painting is an AI video generation platform launched by Baidu, based on AI technology to help users easily create personalized videos. Painting intuitive interface, powerful tools, with inspiration recommendation function, can provide creators with creative inspiration, support a key to the same operation, can quickly generate similar videos, simplify the creative process.
1yrs ago
066.4K
袋鼠参谋 – 美团推出的商家AI智能决策应用

Kangaroo Staff - AI Intelligent Decision Making App for Merchants by Meituan

Kangaroo Staff is a merchant-oriented AI intelligent decision-making application launched by Meituan to help merchants solve problems in store opening and operation. Based on Meituan's massive catering data and more than 10 years of online operation experience, through conversational interaction, it provides merchants with precise information on track selection, store opening location, dish development, store operation and other scenarios, such as...
1yrs ago
066.1K
11ai - ElevenLabs推出个人AI语音助理

11ai - ElevenLabs Launches Personal AI Voice Assistant

11ai is an AI voice assistant launched by ElevenLabs, with voice interaction as the core, through natural and smooth dialogue to enhance the user's work efficiency. 11ai supports more than 5,000 voices, and users can customize the exclusive voice, the assistant is more personalized. With low-latency voice inter...
1yrs ago
066.1K
CombatVLA - 淘天集团推出的高效VLA模型

CombatVLA - Efficient VLA Model by Amoy Group

CombatVLA is an innovative 3D action role-playing game (ARPG)-specific model from the Future Life Lab team of the Amoy Sky Group.CombatVLA is a visual-linguistic-action (VLA) model, built on a 3B parametric scale, that collects human player's through a motion tracker...
11mos ago
065.6K
MineContext - 字节开源的主动式上下文感知AI伙伴

MineContext - Bytes Open Source Active Context-Aware AI Partner

MineContext is an active context-aware AI partner open-sourced by the ByteDance Viking team to help users efficiently manage massive amounts of information and improve the efficiency of knowledge work. Over the screenshot and content understanding technology, automatically record the user's daily operations (such as browsing the web, editing documents, etc.), support...
10mos ago
065.5K
FastVLM - 苹果公司推出的视觉语言模型

FastVLM - Visual Language Model from Apple

FastVLM (Fast Vision Language Model) is an efficient visual language model introduced by Apple Inc. With FastViTHD hybrid visual coder as the core, it incorporates convolutional and Transformer architectures to significantly reduce visual...
11mos ago
065.2K
DeepSeek V3.1 - DeepSeek推出的最新开源AI模型

DeepSeek V3.1 - Latest Open Source AI Models from DeepSeek

DeepSeek V3.1 is a new generation of AI models introduced by DeepSeek, with important upgrades based on its predecessor, V3. DeepSeek V3.1 introduces a hybrid reasoning architecture that allows the model to flexibly switch between thinking and non-thinking modes, significantly improving the thinking...
11mos ago
064.8K
靠岸妙写 - AI论文写作工具,构思到成稿一站式解决

Leaning Wonderful Writer - AI essay writing tool, one-stop solution from idea to finished paper

Leaning Wonderful Writer is an AI dissertation writing tool that provides an efficient and convenient solution for academic writing. The tool supports one-click generation of dissertation outlines, abstracts and first drafts of the body of the paper, which is applicable to different levels of academic needs such as undergraduate and master's degrees, and covers multi-disciplinary fields such as science and technology, liberal arts and social sciences.
1yrs ago
064.6K
Goedel-Prover-V2 - 普林斯顿联合清华和英伟达等开源的定理证明模型

Goedel-Prover-V2 - Princeton's open-source theorem proving model in conjunction with Tsinghua and NVIDIA, among others

Goedel-Prover-V2 is an open-source theorem proving model jointly released by leading organizations such as Princeton University, Tsinghua University, and NVIDIA. The model is based on innovative techniques such as hierarchical data synthesis, verifier-guided self-correction, and model averaging to significantly improve the performance of automated formal proofs...
1yrs ago
064.3K
ChatGPT Agent – OpenAI推出的通用智能AI Agent

ChatGPT Agent - General Intelligence AI Agent by OpenAI

ChatGPT Agent is a general-purpose AI Agent from OpenAI that combines multiple capabilities to autonomously accomplish complex tasks. Users only need to describe their needs in natural language, and the Agent can automatically select the appropriate tools, such as browsing the web, extracting information, running code...
1yrs ago
064.2K
羚珑 - 京东推出的AI商品图设计工具

Antelope - AI product image design tool launched by Jingdong

Antelope is an intelligent design tool launched by Jingdong, providing efficient and convenient design solutions for e-commerce merchants and individuals. Through intelligent keying, intelligent layout, intelligent color matching and other functions, it helps users to quickly generate high-quality design works to meet the main picture of the product, advertising Banner, store page and other e-commerce store...
1yrs ago
064K
Seed GR-3 - 字节跳动Seed团队推出的通用机器人模型

Seed GR-3 - Generalized Robotics Model from the Wordpress Seed Team

Seed GR-3 is a general-purpose robot model introduced by ByteDance with strong generalization ability to adapt to new environments and complex commands. The model fuses visual, verbal, and motion information, and is based on a three-in-one training method of robot data, VR human trajectory data, and publicly available graphic data to enhance the ability to respond to new objects...
1yrs ago
064K
ChatFlow - 开源AI工作流自动化工具

ChatFlow - Open Source AI Workflow Automation Tool

ChatFlow is an open source AI workflow automation tool that supports the transformation of complex requirements into efficient workflows. Tools based on AI technology to help users quickly generate code frameworks, test cases, can assist in writing and designing software architecture.
1yrs ago
063.8K
Qwen-Image-Edit - 阿里通义开源的图像编辑模型

Qwen-Image-Edit - Ali Tongyi open source image editing model

Qwen-Image-Edit is an all-purpose image editing model introduced by Ali Tongyi, built on the Qwen-Image architecture with 20 billion parameters. The model combines both semantic and appearance editing capabilities, and can perform low-level visual appearance editing on images (e.g., adding, deleting...
12mos ago
063.7K
有道小P - 网易有道推出的新一代AI全科学习助手

Youdao Xiao P - A new generation of AI general learning assistant launched by NetEase Youdao

Youdao Little P is an AI all-subject learning assistant launched by NetEase Youdao, designed for K12 students, equipped with the Youdao Ziyi education big model, covering elementary school, middle school and high school all-subject Q&A, providing personalized learning advice. With AI word search and AI translation functions, Youdao Little P helps students quickly solve language problems...
1yrs ago
063.5K
JoyHallo - 京东开源的AI数字人模型

JoyHallo - Jingdong's open source AI digital human model

JoyHallo is an open source AI digital human model from Jingdong, designed for Mandarin, supporting the conversion of audio into realistic speaking video.JoyHallo embeds audio features based on the wav2vec2 model, using a semi-decoupled structure to improve the accuracy of lip movement prediction, and supports the generation of English video...
1yrs ago
063.4K
InternVLA·N1 - 上海AI Lab开源的端到端双系统导航大模型

InternVLA-N1 - Shanghai AI Lab Open Source End-to-End Dual System Navigation Large Model

InternVLA-N1 is an open source end-to-end dual-system navigation macromodel from Shanghai Artificial Intelligence Laboratory. Using a dual-system architecture, System 2 is responsible for understanding linguistic commands and planning long-range paths, while System 1 focuses on high-frequency response and agile obstacle avoidance. The model is trained entirely based on synthetic data through large-scale digital ...
11mos ago
063.4K
VoxCPM - 面壁智能联合清华开源的端到端TTS模型

VoxCPM - Faceted Intelligence and Tsinghua Open Source End-to-End TTS Model

VoxCPM is a speech generation model jointly open-sourced by Facade Intelligence and Shenzhen International Graduate School of Tsinghua University.VoxCPM adopts an end-to-end diffusion autoregressive architecture to generate continuous speech representations directly from text, breaking through the limitations of traditional discrete disambiguation. Through hierarchical language modeling and finite state quantization...
10mos ago
063.4K
DeckSpeed - AI PPT制作工具,自然语言生成演示文稿

DeckSpeed - AI PPT Maker, Natural Language Generated Presentation

DeckSpeed is an AI presentation creation tool based on conversational interaction, where users express their needs based on natural language and quickly generate personalized slides without relying on traditional templates. The tool supports real-time feedback adjustment, users can modify the color, style and content of the slide at any time to ensure that the presentation is complete...
1yrs ago
063.3K
InternVLA-A1 - 上海AI Lab开源一体化操作能力的具身大模型

InternVLA-A1 - Shanghai AI Lab Open Source Integration of Operational Capabilities for Embodied Large Models

InternVLA-A1 is a large model of embodied operation open-sourced by Shanghai Artificial Intelligence Laboratory. It has the ability to understand, imagine, and execute the integration, and can accurately complete the task. The model fuses real and simulated operational data, and automates the construction of massive multimodal through large-scale virtual-real hybrid scene assets...
10mos ago
063.2K
gpt-realtime - OpenAI最新推出的AI语音模型

gpt-realtime - OpenAI's newest AI speech model

gpt-realtime is an advanced speech model from OpenAI that supports direct audio processing to generate natural and smooth speech. The model supports multiple languages and styles, understands non-verbal cues such as laughter, and can switch between languages.
11mos ago
063.2K
Genie 3 - 谷歌推出的通用世界模型

Genie 3 - A Universal World Model from Google

Genie 3 is a next-generation universal world model from Google DeepMind that enables the generation of highly dynamic and coherent virtual worlds in real time.Genie 3 simulates physical phenomena, natural ecosystems, and supports the creation of fantasy and historical scenarios. With text prompts, users can...
12mos ago
063.1K
琴乐大模型 - 腾讯推出的AI音乐创作模型

Piano Music Big Model - AI Music Composition Model by Tencent

Qin Music Grand Model is an advanced AI music creation grand model jointly launched by Tencent AI Lab and Tencent TME Tianqin Lab. The model intelligently generates high-quality stereo audio or multi-track sheet music based on user-inputted keywords, descriptive statements or audio clips in English and Chinese.
1yrs ago
063K
RoboOS 2.0 - 智谱开源的跨本体具身大小脑协作框架

RoboOS 2.0 - Wisdom Spectrum's Open Source Cross-Ontology Embodied Brain-Size Collaboration Framework

RoboOS 2.0 is an open-source framework for cross ontology brain collaboration, promoting the transformation of robots from single intelligence to collaborative group intelligence. The framework realizes efficient division of labor with a "big brain" architecture, where the cloud brain is responsible for complex decision-making and collaboration, and the small brain module focuses on executing specific skills.
1yrs ago
063K
SkyReels-A3 - 昆仑万维推出的音频驱动数字人创作工具

SkyReels-A3 - Audio-Driven Digital Human Creation Tool from KunlunWangwei

SkyReels-A3 is an audio-driven digital human creation tool from Kunlun World Wide Group. SkyReels-A3 is an audio-driven digital human creation tool, which can generate high-quality dynamic video content through simple inputs (e.g., portrait images and voice), make static photos "come alive", and replace lines for existing videos with new lip-syncs that the characters will automatically...
12mos ago
062.7K
ThinkSound - 阿里通义推出的音频生成模型

ThinkSound - Audio Generation Model launched by Ali Tongyi

ThinkSound is the first CoT (Chain Thinking) audio generation model introduced by Ali Tongyi's speech team. The model can generate accurately matched sound effects for video images, based on the introduction of CoT reasoning, to solve the problem of traditional technology is difficult to capture the dynamic details of the screen and spatial relationships.
1yrs ago
062.7K
MagicTryOn - 浙大和vivo等机构推出的视频虚拟试穿框架

MagicTryOn - Video Virtual Try-On Framework from ZJU and Vivo and others

MagicTryOn is an advanced video virtual try-on framework launched by the School of Computer Science and Technology of Zhejiang University in collaboration with vivo and other organizations. The framework replaces the traditional U-Net architecture with an innovative Diffusion Transformer (DiT) architecture, combined with a fully self-attentive machine...
1yrs ago
062.6K