Half Year 2026 MiniMax Group Inc Earnings Call

Speaker #2: 大家好,欢迎参加 MiniMax 2026 年半年度业绩交流会。下面把时间交给公司投资者关系总监 Olivia Sha,请发言。

Speaker #3: 谢谢各位投资者朋友,大家晚上好,欢迎参加 MiniMax 2026 年中期业绩电话会。出席今天电话会议的管理层包括 MiniMax CEO 严俊杰、COO 运业一,以及副总裁 薛子钊。下面我先简要宣读免责声明:今天的讨论可能包含前瞻性陈述,包括但不限于对公司业务前景及预期财务表现的展望。此类陈述受多项风险和不确定因素的影响,实际结果可能与前瞻性陈述所表达或暗示的内容存在重大差异。本次电话会议将主要以中文进行,包括管理层发言和问答环节,并由第三方口译团队提供英文同声传译。英文翻译仅为方便与会者理解,如英文翻译与管理层的中文原始陈述存在任何差异,请以中文原始陈述为准。下面我把时间交给 CEO 严俊杰。

Speaker #4: 各位投资者、分析师,大家晚上好。感谢大家参加 2026 年的中期业绩会。今年以来,我们陆续升级了 M2 系列的模型,推出了 M3 和 A3 等模型。随着模型智能水平的提升,以及推理的吞吐的增加,业务上我们整体实现了快速的增长。7 月份的投控消耗量相比于今年 1 月份基本上是增长了 20 倍,然后 Q3 的收入相比于 Q1 环比增长了 81.8%。在 8 月份,AR 进一步提升到了超过 $800 million。随着 M 系列和 H 系列两个模型管线的持续完善,公司预计在未来一段时间内,模型发布的周期会显著缩短,同时智能提升的速度会变得更快。今天借助这个机会,我来分享一下我们的定位和战略。 首先是大语言模型带动的智能提升,目前几乎没有尽头。从去年下半年开始的 coding 和 edit 能力的使用,到后面更加自主性和创造性的长程任务,再到未来可预期的能够端到端交付结果,我们持续在追求更高等级的智能。但是算力对每家公司来说都是有限的,通过 minimize inference cost 和 maximize intelligence 来实现 intelligence with everyone,这是我们一直以来坚持的技术路线。我们认为这不只是提供性价比的服务,更是面向更高级的智能进行有效 scaling 的核心能力。原因是后续的规模在模型能力中的作用越来越大,而 inference 的成本决定了后续的规模。对于模型来说,首先看智能等级,其次看性价比。模型的智能等级跟模型参数规模有直接关系,而我们的选择是在模型规模上持续研发行业最大的档位,同时我们探索自由的架构来实现预先的 loss 和计算效率的平衡。再借助计算效率优势在后续内容中尽可能 scaling。 首先是预训练,在大参数模型的架构选择上,我们分别探索两种极限。第一个是混合专家的稀疏,我们专家的稀疏度在 2% 以内,依然能够保持良好的收敛性。第二个是 attention 的稀疏,我们提出了 min-max slash attention 的架构,在保证训练稳定的同时,能够大幅压缩长上下文的计算效率。在我们正在研发的旗舰模型中,我们进一步提出了 minimize slash attention 2.0 版本,进一步压缩了 KV cache 需要的存储容量,使得在推理时能够增加 profiling 的缓存利用率,以及增加 batch size,从而使算力利用率变得更高。在最大参数量的这个档位上,我们的计算效率相比于上一代 MSA 1.0 进一步提升了 3 倍,并且在越长的任务上优势越大。 其次,在 middle training 和 post training 上,算力占比实际上是越来越多的。训练 token 有越来越多来自于合成数据、强化学习训练过程中的 rollout,以及支撑实验的大量评测。这三种计算模式实质上都是推理,尤其是随着长任务占比越来越高,长上下文下的推理效率就变成了 scaling 的关键。推理越高效的结构,越容易在 middle training 和 post training 中 scaling 得更快,从而长期占据优势。在相同的算力规模下,可以认为推理效率提升的倍数基本上跟后续规模变大的倍数显著相关。我们在大参数模型架构上带来的优势和创新,以及后训练规模的持续扩大,预计会在后续模型更新中有非常直接的体现。 接下来讨论一下技术设施。在加强与云厂商合作的基础上,为了进一步增加算力供给,我们实际上是国内最早的两家建立规模化、长期稳定的技术设施的独立大模型公司之一。这是我们能够同时研发两款模型的关键支撑。目前自建的技术设施已经能独立支撑我们目前最大的训练任务,其 ETTR(有效训练时间比例)现在能达到 97%,是行业内目前所知最高的一档水平。这种全栈能力不仅能够提升训练规模,还使我们在体系结构设计上有很高的灵活性,从而在 AI 原生的训练和推理上有大幅改进的空间。比如,借助内存、SSD 和网络的协同设计,可以实现 KV cache 利用率的显著提升,成倍降低 profiling 的成本。随着后训练中 edit 类任务占比提高,需要大量以 CPU 为基础的沙盒,而我们多模态业务中的推理集群,有大量 CPU 的混合调度空间。同时,线上的推理流量潮汐性质,也给我们做小规模验证实验和强化学习 rollout 提供了大量可复用的算力。基于我们在 infra 上的全栈能力,单位资金效率下后续规模,预计至少还有 3 倍的空间,并且推理效率的提升能持续增加后续的毛利水平。 最后讨论多模态。跟大部分大模型公司不同,我们在设计模型时一直原生考虑多模态信息。我们认为世界的理解和生成是生产力的重要组成部分。多模态生成目前也是 AGI 内除编程外的第二大市场。H3 实际上是语言模型和视觉生成相结合的实践。这不仅是把语言模型作为 encoder,更重要的是利用语言模型做精细的上下文感知,实现全方位的参考和精细的生成能力。min-max H3 的效果突破,以及我们选择的开源策略和展现出来的性价比优势,实际上改变了该领域的先进模型基本都属于大厂且走封闭路线的格局。在创作者和开源社区中的受欢迎程度远超预期,自一三开源到现在三周多,min-max H3 的下载量已超过 2,400 万次,诞生了 300 多个公开衍生模型,这应该是今年全球下载量最大的模型。借助良好的社区口碑和性价比,我们线上官方服务的使用量也有爆发式增长。 随着多模态和语言模型的深度融合、物理世界信息的引入、encoder 上下文计算和数据等环节的 scaling,我们认为接下来的多模态模型还会有多次能力边界的跃升。基于我们在语言模型和视觉模型两方面的技术储备,我们认为后续在该领域的技术领先性、市场份额和影响力也会进一步增加。 更长远来看,我们认为类似语言模型和视觉模型结合的范式,在接下来一段时间会在一些更前沿领域持续涌现。回到最开始的问题,min-max 追求的不只是智能和成本之间的取舍,只有最小化单位致能成本,才有可能进一步训练出最大化的智能水平。只有降低单位智能供给成本,才可能让更高水平的智能进入更广泛的生产和生活。 模型创新决定我们能把智能推动到多高,而模型效率决定智能能提供多大供给。强化学习和现实世界反馈决定下一轮智能进步的速度,而多模态和专业领域决定智能能进入多深多广阔的真实世界。我们也清楚单个模型并不代表全部,核心还是背后研发模型的基建和研发体系能否持续 scaling。 在 M3 的研发过程中,我们确实曾有时间上的不足,但对我们而言,更重要的是能否持续定义和迭代自身技术路线,能否不断加速智力提升的密度,能否将更高水平的智能以更低成本交付给更多用户。 这也是我们公司坚持的愿景:intelligence with everyone。感谢大家的关注和支持,我们会继续努力,谢谢大家。

Speaker #1: 下面有请公司副总裁薛子钊介绍财务情况。

Speaker #3: 大家好,我是子钊。下面我简要介绍一下集团 2026 年上半年的财务表现。报告期内我们实现了收入约 1.2 亿美元,同比增长 283%。上半年的收入已达去年全年收入的 1.5 倍,其中第二季度收入相比于第一季度环比增长 82%。收入的增长的核心在于我们围绕文本及多模态模型形成了独特的定位,在提供前沿模型能力的同时,优化推理效率,实现了具有吸引力的性价比。凭借这一优势,我们的全球企业客户和开发者规模快速扩大,API 调用量及模型推理需求持续增长,推动模型能力加速转化为实际产品使用和商业化收益。从收入结构上看,开放平台及其他 AI 企业服务收入约 7400 万美元,同比增长 703%,占总收入的比例由上年同期的 30% 提升至 63%。主要由付费用户及企业客户数增长,API 调用量增长,以及 token plan 快速获得采用所带动。AI 原生产品的收入约 4300 万美元,同比增长 101%。海外市场贡献了超过六成的收入。进入第三季度,收入加速度增长,目前 8 月 ARR 已超过 8 亿美元,其中 2B 板块在整体 ARR 中的贡献已超过 80%。在技术和基础设施层面,随着 min max sparse attention 等模型架构创新落地,以及基础设施能力的持续增强,我们不断提升模型的计算效率。2026 年上半年毛利约 2100 万美元,同比增长 465%。毛利率有 12.1% 提升 5.8 个百分点至 17.9%。费用方面,研发开支约 3 亿美元,同比增长 139%。主要由于我们持续投入基础模型能力的研发,研发开支增速显著低于收入增速,体现研发成果转化为业务增长的效率进一步提升。销售及分销开支约 0.3 亿美元,同比下降 18%。主要由于用户自然增长的策略持续推进,推广开支相应减少。行政开支约 0.3 亿美元,同比增长 104%,占收入比例由 49% 下降至 26%。经营杠杆逐步体现。亏损方面,IFRS 会计口径期内净亏损由上年同期约 4 亿美元,收窄 11% 至约 3.6 亿美元。经加回股份支付,金融负债供应价值亏损及上市费用后,2026 年上半年经调整净亏损约 2.9 亿美元。上年同期约 1.4 亿美元,同比增加 111%。主要反映了我们对于基础模型能力和基础设施的持续投入。为进一步增强公司的资金储备和财务灵活性,2026 年 7 月公司完成了配售发行合计款项总额为 160 亿港元,目前公司现金储备超过 30 亿美元。为我们持续投入前沿模型研发,算力基础设施建设及吸引顶尖人才提供了坚实的保障。总体来看,上半年公司的收入规模、业务结构和毛利率均取得明显进展,费用指标持续改善。我们将继续在保持长期战略投入的同时,提升研发、运营和商业化效率。谢谢大家。

Speaker #1: 谢谢大家。以上是公司的发言内容。接下来可以进行 Q&A 的交流环节。

Speaker #2: 感谢领导的分享。下面进入互动交流环节。欢迎各位投资者举手语音互动或发送文字提问。

Speaker #4: 大家好,如需提问,电话端的参会者请先画机上的星号键,再按数字 1。网络端的参会者,您可以在直播间互动区域内文字提问,或点击旁边的"举手"按钮申请语音提问。谢谢。大家好,如需提问,电话端的参会者请先画机上的星号键,再按数字 1。网络端的参会者,您可以在直播间互动区域内文字提问,或点击旁边的"举手"按钮申请语音提问。谢谢。

Speaker #2: 下面有请中信证券杨泽源进行提问。有请。

Speaker #5: 好的,感谢各位 MiniMax 馆长给我这个提问的机会,也非常恭喜公司上半年业绩取得了这个长足的快速的进展。我想请教两个问题。有一块说了,第一个问题是这样,就是市场其实最近比较关注这个 coding 的这个竞争的加剧,当然也包括部分的其实已经在海外的模型商业化增速的变化。刚刚其实 IO 讲的过程中,我们也在认真的学习。就是我看咱们提到说,咱们判断全球的这个模型智能仍然有很大的提升的这个空间,这个是为什么做出这样的判断。然后进一步讲说,在 coding 之后,哪些的,比如说能力的突破,或者说应用场景,最可能成为下一轮这个规模化需求的来源。这是第一个问题,就是关于 coding coding 之后的这些事的判断。第一个事,我一块问了第二个问题,谢谢。第二个问题是这样,就是公司刚刚也披露了 7 月份的这个 token 的消耗量,其实也达到了这个应该 1 月份,我如果没记错应该达到 1 月份的 20 倍,而且 8 月份的这个 ARR 应该也增长很快,应该也达到,进一步超过了 8 亿美金,对吧?那能否辛苦管理层具体去拆解一下这个增长,它主要是来自于是新增客户,还是说是现有客户的扩容,还是说是模型能力的提升,或者说是我们因为我们有动态,有文本。还是来自于这个产品结构的变化,这个来源是什么?然后现在咱们 BC 的这个结构,或者说地域的结构,就是区域的这个结构大概是怎么样,也就是我第二个问题就辛苦领导能不能帮我们拆解一下咱们刚刚讲到的这些增长的数据。谢谢,就这两个问题。 好,这个感谢这个提问。首先第一个是关于 coding 相关的,我觉得这个智能提升空间其实它既有这种技术上的这种变化,然后也有这种来自市场上的变化。首先我们先说技术上,就技术上的话,实际上在过去几个月,实际上的话,就是模型智能提升的一个关键的评价指标其实是说衡量模型能够做复杂长任务的这种能力。那在这个就长任务的话,指的是说比如说一个任务,它经常比如说需要做这种几百轮的这样的这个就是需要做几百轮的这种比如说工具调用,然后比如说 coding,比如说结果验证,然后类似这样,就是类似这样级别的一个复杂的任务。那从过去几个月外面的进展,包括我们内部能够做的大量实验,我们可以观测到一个趋势,就是说只要是一个能力,它可以被精确的评估,然后它可以被构造出来这种强化学的这种学习的环境,基本上就是基本上就是说模型基本上就可以学出来,这个是一个我觉得这已经是一个非常确定性的一件事。那这里面的这个核心的竞争力或者是核心的改进方向其实是来自于是说怎么样把这些专业领域的问题把它抽象成一个一个这种高质量的这种强化学的环境,然后就是让可验证,让模型能够持续来学习,然后就是基本上就是说比如说模型来提升这些能力,那在一个足够大的一个预训练的基础之上,接下来就来自于比如说 SMT 和强化学习,基本上就是在 SMT 的时候,它比如说这个里面的话,就是说数据的学习效率相对偏低,所以可以用比较大量的数据,然后在强化学习的时候,基本上就是用这种最高质量的数据,然后实际上我觉得就是根据我们刚才提到的,就是说我觉得在过去跟接下来一段时间,我觉得这个强化学习的这个环境和规模可能会区分不同的公司的这个模型进步的速度。然后实际上我们已经投入了大量的精力来把善这件事情,然后那个从这个市场的角度来说,就是说我们认为就是说首先 coding 其实是所有的能力里面目前可能是最核心的一个能力,那其次的话是说这个但是即使这样,我们发现就是说单纯的 coding 的能力其实还是有很多的提升空间的,比如说其实很显然就是现在 coding 的使用还是停在代码生成和一些辅助开发的阶段,然后这些更加端到端的这种交互的这样的任务里面,现在其实我们认为提升空间还是非常大的,然后这是第一个。然后第二个的话是说我觉得现在其实出现了越来越多的这种就是泛 coding 类的这样的任务,比如说举例子,比如说像网络安全,在过去一段时间大家看到的其实就是一个非常典型的例子,就是从这本质上就是说当一个模型的 coding 能力足够强,然后再加上大量的这种安全相关的环境和迭代之后,那在网络安全上其实展现出来这种惊人的能力,甚至说改变了这个网络安全这个行业。那我们现在其实能看到是说就是只要是说能够建立这种数据的闭环,能够这种可被验证的来构造出来这样的强化学的环境,其实很多行业都能看到这样一个趋势,比如说像芯片设计,比如说一些类似的这样一些领域,我觉得大家可以预期的事情是说可能在接下来几个月这样的事情可能会出现好几次,这是第一个。然后第二个的话是说就是我想说一下这个就是商业化相关的,就是说其实可以发现是说当目前能力通过一个边界之后,就是模型的这个增长,不管是行业还是我们公司,它其实不是这种连续的,它其实是说一个好的模型出来之后,它其实会有一些这种阶梯式的这样的释放,而不是简单的这样一个这种连续的线性的变化。我们其实在这个今年7月份,就是说我们的这个 token 消耗量的相比月份多了很多,它其实主要是来自于两步的增长,第一步的话是在今年春节的时候,就是一些这种 agency 的这种任务上,它其实取得了一个非常大量的增长,然后第二个的话是说在 M3 出来之后,我们其实提供了这种原生的动态能力,coding 能力其实也有一定的加强,然后同时的话和计算速度其实也变得更快,其实就进一步带动了这种 token 消耗量的一个增加。对,然后那就是说接下来讲一下第二个事情,就是说关于这些这个收入数据的一个拆分,首先就是说8月份的这个 AR 超过8亿美金,其实这个核心的驱动力还是来自于。模型能力的提升,我们其实一直在践行这种极致的智能的极致性价比,在 M3 在相比于 M2 的这个模型价格不变的情况下,其实我们首先提供了更高的智能水平,然后第二的话是说我们其实吸引了不少新的客户,它的这个企业客户和发展的数量,我们前段时间统计了一下,其实已经超过了200万,其实是去年年底,其实是去年年底的10倍左右。然后另外一方面就是说老用户,老客户其实也是因为模型能力的提升而进一步来解锁了更多的场景,比如说举例子,比如说除了 coding 跟 agent 之外,比如说很多办公的场景,比如说很多客户会进一步来使用这种办公的场景,我们观测到就是说这种模型的消费正在从这种人类与 AI 的这种交互变成了这种 agent 与 AI 的这种多元的交互,然后人类的问题,然后通过 agent 的形式被展开成这种多轮的这种模型请求、工具调用和 agent 任务,而这种由于这种 agent 驱动的这种推理需求的增长,其实显著快于这种人类用户的数和消息数,然后这些变化其实也发生在我们的模型身上。从结构维度来看,就是说这种单个用户的这种 token 消耗其实是在快速的增长。刚才还提到这个 AR 的这个结构,这边的话是这样的,就是说从客户结构来看,就是说在当前就是在这个8月份的时候,这个里面的这个 to be 的这一部分,它的占比应该是80%左右,然后在去年的时候,它的这个 to be 的占比其实是大概是30%,然后在这个里面的话,我们可以看到这种企业和开发者对模型的采用速度其实是在明显加快的,我们认为就是说在这个AI时代,其实模型本身就是产品,所以不管是to be还是to see,本质上都还是客户为这种模型能力而付费,这个的话也是我们作为一个大模型公司,就是把这种持续提升模型的智能水平作为一个最核心的一个目标,并且不断打磨出来这个更优秀的模型能够被更广泛地用来使用。然后从这个地区结构上来看,我们这个国际化依然是公司非常重要的特点,在这个今年的上半年,我们海外的收入占比其实是超过了应该是60%左右,然后从这个模态的角度来看,就是说我们其实能够观测,当然比如说在A3出来之后,动态的这个增长其实是有了一个非常大的增长,我们其实但是其实远模型其实也依然在增长,本质的原因是因为我们其实提供了非常优秀的这种性价比。然后在这种我们观测到我们的这个客户跟用户,我们可以发现是说这种越来越多的客户会同时使用这种语言、视觉和这种声音的能力,所以说就是说在这个从6月份到8月份,我们认为它的这个其实是语言模型和动态模型是两个都可以加在一起来驱动这种AR的增长。然后如果再往后展望的话,我们认为下一阶段最重要的增长来源还是这种模型能力的变化了,就是更强的模型肯定会带来更多的用户和任务。而且这种更高的推理效率其实也会让客户可以以更加可接受的成本能够扩大使用的范围,我们后续可以期待这种我们接下来几个新的系列,包括M3.1,包括M3 Pro,还有那个H3.1的这个更新。

Speaker #1: 好的,感谢,也期待公司更多的新意,谢谢。

Speaker #2: 好的,感谢领导的解答。下面有请高盛Ronald Kong进行发言,有请。

Speaker #3: 谢谢严总Timo的招呼,Olivia,那两个提问,第一刚才管理层你们提到这个minimize cost,然后maximize intelligence,那我们看到市场上有一些友商是不断地提价的,也有一些模型是在降价的,那公司能否进一步检视一下,给我们说一下那个为什么降低推理成本,那不只是一个商业上的性价比策略,但也会直接影响这个模型智能的上限,就是刚才说的这个maximize intelligence,那顺着这个问第二个提问,那刚才提到M3.1和M3 Pro,那能跟我们分享目前研发和这个发布的节奏如何吗?在coding agent和一些长程的任务上,市场应该关注哪些可验证的进展?谢谢。

Speaker #1: 首先这个第一个问题,就是说其实我觉得推理它其实不光是一个往外来提供这种智能的这么一个服务,它其实也是我们内部能够产生下一代智能的一个重要的一个过程,原话是因为我们可以非常清晰地看到是说一方面比如说在预训练阶段,还是可以跟那个scaling从比如说接下来3T的模型到再往后比如说更大规模的模型,它的预训练的规模其实也是在参数规模和对应的这种token量其实基本上是1比20,还是可以持续能够往scaling的,但是在另外一方面就是说其实后训练的这个占比其实是越来越高的,后训练的话这个里面的话其实是需要依赖于非常多的这种比如说这种数据的合成,比如说依赖于这种比如说agent的这种环境的构造和这种采样,比如说依赖于比如说在强化学习的时候,然后需要做大量的这种rollout,那这些当然了最重要的事情是说因为后训练里面有非常多的不确定性,需要做大量的实验,然后这些实验的话它其实也都需要来做评测,这种非常精确的评测,那在这个里面的话其实可以看到大部分计算量实际上已经是就是在后训练中其实大部分计算量其实是推理,那也是因为这个原因我认为其实推理效率其实是决定了在后训练中我们到底能够scaling到一个什么样的规模。那就是这个的话是就是我觉得基本上可以认为是说这种在同样的规模的模型下推理效率越高,其实它能够从后训练中能够享受到红利其实就越大。然后推理效率的核心其实它也不光是说这种单个token的成本,其实是说在单位时间相同的算力下能够完成多少次的这种有效推理,然后这种更多的推理就意味着更多的训练轨迹,更多的实验,然后更快的评测,所以说它的迭代其实就会越来越快。那这个的话是我们就是为什么这么重视推理效率的原因,就是从这种对外服务的成本和对内的这种模型的研发的加速和scaling上,我觉得它基本上都是这两方面其实都非常的重要,这个的话也是我们认为是说我们下一代模型的这个非常核心的竞争力,然后以及是说是我们再往后的模型上,在性能上能够取得领先的基础。然后就是我们说一下这个接下来的这个模型更新的节奏,实际上我觉得模型的发布它其实只是一方面,它更实际的事情是说在研发的模型的过程中,能够能不能形成这些可以沉淀、可以积累的这样一些基建,然后从这个数据到算力到预训练的这种框架、强化学习的框架、评测,然后包括这种推理的基建,然后当然对我们来说还有更底层的,包括这种规模化的集群建设、软件的优化,然后再到这是整个环节,很多这种基建和研发同事、算法同学、优化同学之间的这种相互的配合,这其实是一个非常完善的过程,实际上的话就是说我们之前在那个M3的时候,其实我们经历过一些这种比较有挑战的时刻,它的背后其实反映的事情是说我们的这个评测的基建其实是没有做好,那。 We have been advancing fairly well on these products. For M3.1, the positioning is to smooth out, or build a closed loop of our underlying infrastructure.

And also, we have been trying to raise the capability of the H3 a little bit higher for the next generation.

Okay, thank you.

Uh, just a quick add

for all of these models.

The key is to make sure that we do every

thing. Well,

making sure that we have a robust infrastructure.

Making sure that we introduced and made these models accessible to users as soon as possible.

Thank you.

Next question, from Gary at Morgan Stanley. Please go ahead.

Thank you, everyone.

My question is on compute.

Which is a key resource for large language model companies?

The country is quickly ramping up the production of, uh, chips. How about our compute reserved for our M3 and also their future models? Do we have sufficient compute resources? Also, what's the breakdown for training and inference, and the, uh, domestic chip compatibility? That's something I want to know as well. Thank you.

Well, that's a key question.

Overall, our

current computer or planned computer resources, sufficient to support the iteration of our planned, uh,

Products pipeline.

And also, thanks to our in-house build-out experience.

We are actually planning or building the reserve for building more, uh, computer resources for the future. Uh,

Interactions.

For models of 3 trillion parameters.

As well as the Token Factory, we have built a comprehensive compute network.

For ecosystem.

so, with our internal coordination,

So we achieve—so we secure computer resources from our domestic or self in-house, to compute computer resources from hyperscalers and also Token Factory.

So our self- or in-house computer that is used for the flagship model training is for large-scale computer clusters.

I mean, it runs consecutively, and it has high requirements for both software and hardware coordination efficiency.

So, we need to

build these auxiliary.

High-bandwidth network and dedicated, uh, network.

Etc, and Etc.

So, these will be reflected in our, uh, financials.

however, I

do want to say that we have achieved notable progress in the first half of the Year and that lays the foundation for our quick future product Integrations. In moving, on to the deepening collaboration, with the cloud, uh, Pro providers, which provided sufficient, uh, Supply computer supply, or flexibility for our training,

And also influence, uh, workload.

Our actual inference workload is actually dependent on.

Or directed to the, uh, hyperscalers.

And also, we are expecting the, uh, network, you know, giving us a lot of options.

And in addition to the two resources, we are also working with the Token Factory or alternative inference workload compute.

So based on the actual token request for demand, we actually—

Increased the supply.

You know, with this, we are able to accommodate the peak demand.

This allows us to provide a quality service while increasing the scale of their business for these token factories.

We have different models under R&D, and the internal coordination or resource coordination.

Is, um, something like this unique to us? I think.

For different models.

At different stages.

The demand is currently changing dynamically.

Text is at a key. Is that a key stage?

So, the training workload is significantly larger than that of the video generation models.

Large language models are also.

Uh, providing a strong foundation for other multimodal products.

You.

We are actually using some domestically produced chips, and in Q4,

The share of domestic chip is

will increase.

Uh, from, you know, across different models, we observed that the domestic...

Supply chips will

reduce the, uh,

Cost per token.

I mean, it's a very good choice—I mean, for us to secure compute resources.

Okay, thank you.

Thank you from citic.

Thank you very much for the opportunity.

Our gross margin gain was quite significant as a token.

The supply chain or management efforts—whether operational, leveraged, et cetera.

This is something that we value tremendously.

I think the margin gain comes from a combination of efforts, and in the

And there's still significant room for improvement, number one.

Which is the most important thing?

Optimize the use of our, uh, computer resources.

Actually, the throughput of our, uh, unit per computer actually increased three times, actually.

The pricing.

An increase in offset is responsible for some of the gains.

But nevertheless, we have seen significant improvements in that. Number two.

our computers used for training and

inference.

We are coordinating these resources effectively for higher utilization.

For example, during the nighttime,

There's going to be, going to be less, uh, uh, inference.

Requests will then be directed for evaluation and algorithm assessment.

This pushes up the utilization of our computer resources.

Because we are.

in-house infrastructure, so there's greater flexibility, allowing us to achieve

You know, these resource coordination.

Number 3.

Supply chain.

As we?

Increase our in-house build out.

There. So, there's room for improvement from the supply chain management side as we—

See these costs, uh, come down.

We?

Are also.

Uh, optimizing our architectures to achieve better scaling. So, from the tech side,

This is something that's unique, I mean, to us, that we have been exploring.

and for us,

As we reduce the costs, we're going to

Lower the cost of our products to make our products accessible to more users with more supply, and with the scale, we can also improve the margin further.

We don't believe it's a conflict between lowering prices and also improving the margin.

As long as we build a positive, uh, cycle—I mean, the higher scale would provide further room for higher, uh, gross margin improvement. So, in terms of our business mix, our multi-modality, and also audio,

Uh, generation models are having higher margins.

So, the text models are quickly contributing to the revenue growth and also...

As we further optimized.

And improve our computer clusters.

We expected that. The text models will be a key driver in improving our gross margins, and we're looking beyond that.

In the second half of this year, we will continue to see margin improvement, and even further improvement beyond 2026.

Thank you very much.

Next question.

Thank you, management, for the opportunity, and congratulations on the strong ARR in August. I've got a two-part question regarding the HS3 launch and open source.

so how does management uh,

Validate.

The path.

For multimodality and how open source will.

Let’s not undermine our pricing, commercialization, or competitiveness.

Balance the investment.

over the

You know, or a long period of time in the past.

We have been repeatedly asked by many investors.

About whether we pursue the multi-modality strategy.

As the.

I think the H3 reputation definitely, you know, proves—

Our decision—I want to explain why we need to do this open source.

Number 1.

we believe this is a

We're talking about productivity.

It is closely related to

Productivity, you know, when we are

Launching H3, we were aimed.

We are aiming to change the industry because this has been dominated by closed source.

Highly priced, uh, uh, models—I think like language models—we believe a virtuous cycle...

Is 1.

That.

Is characterized by openness.

And that's the reason why we open source our, uh, um, product. Again, we believe that we're still at the very early stage,

So, there's a lot of enthusiasm in the, uh, community.

I mean, it is doing very well, but, uh,

Still, we believed to provide better reliability to generate, uh, uh, better output at a more reliable—uh,

Uh, sources is something that we still need to improve. So we need to keep, uh, push this, uh, industry forward.

With each iteration.

And also, I mean, this is going to increase our capability to monetize these capabilities, and the age of 3 serves as a very good starting point. Moving on to resource,

Allocation.

We?

Do not.

put text models, you know, composition with the

uh, multimodality is actually

Underneath all of these, reinforced by each other, we pursue a coordinated development, creating greater synergy.

And we want to deliver this, uh, productivity tool—like a great productivity tool—to users and make it accessible to more users.

and,

Finally, I want to talk about—uh—what's the relationship between text and video generation models?

How do I view that?

I think going forward.

There are going to be more Frontier domains. I'm thinking of chip design.

Pharmaceutical.

Research.

This is where we can integrate different modalities.

Using inputs from various modalities.

I think we have.

Proven.

A path, which is viable.

How to put it. Um,

so by better understanding,

We?

Can understand how we can increase the synergies through integration.

And the fundamental and ultimately changing the sector, we believe, in Q4.

Or q1 next year.

We believe that.

There will be more emerging opportunities that is going to be, uh, increased integration between the, uh, uh text and uh, video generation or text of versus the multimodality models.

Strong from Jeff.

Please go ahead.

Thank you management.

AIP. So, the competition landscape is intensifying. How does management view the competitive landscape?

And second question.

So what's the core competitive mode? Defining? Uh, Minimax is, uh, competitiveness in the long run.

Thank you. This industry is not a zero-sum game.

Competition.

I mean, by whatever time points, we are just at the starting point.

Looking over the next 3 to 6 months, I mean, the intelligence that we actually serve is still at its initial or early stage.

By that.

I mean.

The key is how fast the speed of iteration is.

In the second half of last year, we have seen the value of models in the use and it's the use of a coding and agent and also agencies are going to be. And also the models will be able to complete a more loan uh Horizon tasks and uh deliver end-to-end capabilities.

So throughout this process, for any company that can keep pushing the frontier of the technology, they can grow the overall market and benefit everyone—or all the stakeholders—in this ecosystem. I think the key in the next phase of competition is not only the

The size of compute.

But rather.

Whether?

The company can Define questions.

Independently, and choose an effective path.

To keep scaling it.

There are, of course, many choices or options in architecture.

Compute.

RL evaluation. There are a lot of options in front of them, so different companies have different attack paths ahead of them. So how do they

Choose.

The right path.

So what, am I going to see, uh, a blossom of four different solutions?

From different, uh, players.

So, as we increase the intelligence,

Delivering the most cost-effectiveness is not only a pricing or business strategy, but also a technological strategy.

For us.

And this is something that we are committed to.

and ultimately,

as AI.

Increases protectivity.

So, we still need to.

Judge or evaluate a model by how well it handles day-to-day, uh, tasks.

And we have seen.

from practices that, uh, users will pay for.

better performing models.

Uh, so we do not need to position the composition, you know, between startups versus big tech companies.

So the key is whether, who can reach the level of intelligence, lower the cost per unit, and increase the conversion efficiency.

So, that is our secret sauce, or our view on your question.

Due to time constraints, we will not be taking any further questions.

So, all of these are based on available information. In case of any discrepancy, please refer to the announcement release on our website.

Okay, that concludes today's call. Have a pleasant evening. Goodbye.

Browse all earnings call transcripts

Half Year 2026 MiniMax Group Inc Earnings Call

Demo
100

MiniMax Group

Earnings

Half Year 2026 MiniMax Group Inc Earnings Call

100

Wednesday, August 26th, 2026 at 12:00 PM

Transcript

No Transcript Available

No transcript data is available for this event yet. Transcripts typically become available shortly after an earnings call ends.

Want AI-powered analysis? Try AllMind →

Earnings analysis guides

Methods for extracting KPIs and checking source support when reviewing an earnings call.

Browse all earnings calls