AI 日报hiw3c.com

梁文锋狙击战:深扒那些梁文锋署名的论文有多牛

雷峰网·AI 科技评论 www.leiphone.com 网页快照

梁文锋狙击战:深扒那些梁文锋署名的论文有多牛

1.《DeepSeek LLM: Scaling Open-Source Language Models with Longtermism》 2.《DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models》 3.《DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model》 4.《DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence》 5.《DeepSeek-V3 Technical Report》 6.《DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning》
7.《Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention》 8.《Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures》 9.《mHC: Manifold-Constrained Hyper-Connections》 10《DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation》 11.《DSec: An Efficient Sandbox Infrastructure for Large-Scale Agent Training》

2024 年初的破局之战:

买不起一万张显卡,那就打效率战

2024 年到 2025 年的能力击穿战:

冲击闭源模型的智力高地

《Native Sparse Attention》《Insights into DeepSeek-V3》

2025 年末的深水区攻坚战:当 AI 冲进无人区,用数学重构底座。

2025 年末的深水区攻坚战:

当 AI 冲进无人区,用数学重构底座

2026 年的终局卡位战:

把战火烧到 Agent 基础设施层

技术狙击手的哲学

《AGI市场观察》发布:四款国产模型周用量破10万亿t ...

深度解读 DeepSeek V4.1 Flash 全新架构,如何成为显 ...

规模要追平智谱!DeepSeek今年将扩招到1000人;继姚 ...

GPT-6 砍掉思考 Token,俄罗斯人砍掉通信 Token,Token 经济开始崩了?

深度解读 DeepSeek V4.1 Flash 全新架构,如何成为显存杀手

打穿 AI 智商测试!GPT-6 Astra 的符号世界模型,是突破还是钞能力刷分?

GPT-6 Sol 降价 50% 的秘密:消失的 Terra,一场模型梯队平移

GPT-6 Sol 降价 50% 的秘密:消失的 Terra,一场模型梯队平移

工作流可以自我进化了,英伟达开源 SoL-Pi,每小时省13.5刀!

深度解读 DeepSeek V4.1 Flash 全新架构,如何成为显存杀手

打穿 AI 智商测试!GPT-6 Astra 的符号世界模型,是突破还是钞能力刷分?