[{"data":1,"prerenderedAt":25},["ShallowReactive",2],{"news:732":3},{"code":4,"message":5,"data":6},200,"操作成功",{"createBy":7,"createTime":8,"updateBy":7,"updateTime":8,"id":9,"title":10,"titleEn":11,"keyword":12,"newsDescribe":13,"urlPath":14,"tourl":15,"articleContent":16,"publishType":17,"briefIntroduction":18,"sort":19,"type":17,"publishStartTime":20,"showTime":15,"publishEndTime":15,"publishStatus":21,"isValid":21,"isOld":19,"remark":15,"nickName":15,"numberOfViews":22,"time":23,"year":24},45,"2026-07-27 09:29:43",732,"WAIC 2026观察：不止于参数，AI竞争迈入「Token产能」时代","732","WAIC、Token产能、Token工厂、统一网关、Token Ops","蓝耘通过统一网关纳管多模型资源，利用AutoModel智能调度引擎识别任务并选择调用路径，同时提供Token计量、成本分析、统一鉴权、调用审计、服务监控和企业级保障。","public/cloud-official/2026-07-27/9b0bab8daab749ef973956e3bf7c97a4.png",null,"\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">在WAIC 2026上，一个明显变化正在发生：AI产业讨论的重点，在关注\u003C/span>\u003Cstrong style=\"font-size: 12px;\">「模型有多少参数」、「集群有多少GPU」的同时\u003C/strong>\u003Cspan style=\"font-size: 12px;\">，开始越来越多地关注\u003C/span>\u003Cstrong style=\"font-size: 12px;\">「每一份算力究竟能生产多少有效Token，又能创造多少业务价值」\u003C/strong>\u003Cspan style=\"font-size: 12px;\">。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">参数依然是衡量模型能力的重要指标，但当大模型进入规模化应用阶段，参数量需要与响应速度、服务稳定性、成本控制等指标共同回答企业更关心的问题——\u003C/span>\u003Cstrong style=\"font-size: 12px;\">Token产能\u003C/strong>\u003Cspan style=\"font-size: 12px;\">。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">AI的竞争，正在从能力展示进入生产运营。\u003C/strong>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Ch3 class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cstrong style=\"color: rgb(46, 120, 255); font-size: 14px;\">WAIC 2026：算力竞争正在增加新尺度\u003C/strong>\u003C/h3>\u003Cp class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">这一变化已经体现在WAIC 2026多家厂商的现场发布中。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cul>\u003Cli style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">新华三推出“图灵Token工厂”，将Token生产、调度、计量与治理纳入全链路运营；\u003C/span>\u003C/li>\u003C/ul>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cul>\u003Cli style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">是石科技发布“国产Token优化工厂”，强调在同等算力投入下生产更多有效Token；\u003C/span>\u003C/li>\u003C/ul>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cul>\u003Cli style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">联想展示以Token驱动的算力基础设施与生产交付体系；\u003C/span>\u003C/li>\u003C/ul>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cul>\u003Cli style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">智微智能则联合生态伙伴发布“Token工厂”。\u003C/span>\u003C/li>\u003C/ul>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">不同厂商的技术路径并不相同，却共同指向一个趋势：\u003C/span>\u003Cstrong style=\"font-size: 12px;\">算力基础设施正在从资源供给走向智能生产。\u003C/strong>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">过去，行业主要关注GPU数量、参数规模和峰值算力；现在，还要衡量Token吞吐、响应时延、服务稳定性和单位Token成本。Token工厂由此成为WAIC 2026频繁出现的关键词，它关注的是如何将算力和模型转化为稳定、可计量的Token产能。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">但“生产出Token”并不等于“运营好Token”。当模型数量、调用规模和业务场景持续增加，\u003C/span>\u003Cstrong style=\"font-size: 12px;\">如何提高有效产出、控制调用成本并保障服务质量，成为企业面临的下一道问题。\u003C/strong>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Ch3 class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cstrong style=\"color: rgb(46, 120, 255); font-size: 14px;\">Token Ops：从统计Token到运营Token\u003C/strong>\u003C/h3>\u003Cp class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">如果把Token工厂看作生产系统，那么\u003C/span>\u003Cstrong style=\"font-size: 12px;\">Token Ops就是围绕这套系统形成的运营方法\u003C/strong>\u003Cspan style=\"font-size: 12px;\">。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">它管理的不只是“调用了多少Token”，而是一次AI任务从输入到交付的完整过程：选择什么模型、消耗多少Token、产生多少成本，以及最终完成了什么任务。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">第一，让消耗可观测、可归因。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">企业需要看清不同部门、客户、应用和任务的 Token 用量、调用成本、响应时延与成功率，为成本分析和策略优化提供依据。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">第二，减少无效消耗。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">在多轮对话、知识库问答和 Agent 任务中，可以通过上下文压缩、信息提炼、输出控制和缓存复用，减少重复信息、过长输出和异常重试带来的额外消耗\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">第三，优化模型路由。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">普通问答和信息提取可以由轻量模型处理，复杂推理和代码生成则调用能力更强的模型。系统还可结合价格、时延、质量和运行状态动态选择调用路径，并在模型限流或服务波动时自动切换，保障业务连续性。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">第四，建立预算与价值评估。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">企业可以设置预算、配额、告警和熔断机制，防止异常调用造成成本失控。同时，将Token成本与任务完成率、质量及业务产出结合，衡量每个有效任务的实际投入。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">Token Ops最终形成的是一个\u003C/span>\u003Cstrong style=\"font-size: 12px;\">“调用—调度—交付—计量—优化”\u003C/strong>\u003Cspan style=\"font-size: 12px;\">的运营闭环。它追求的不是一味压低单个Token的价格，而是在保证质量和稳定性的前提下，用更合适的模型、更少的无效消耗和更可控的成本，完成更多真实业务任务。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">\u003Cimg src=\"https://oss.lanyun.net/public/cloud-official/2026-07-27/34eeb740f524412f9204d9a69a0b4680.png\">\u003C/span>\u003C/p>\u003Ch3 class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 14px; color: rgb(46, 120, 255);\">Token Ops如何进入企业生产系统？\u003C/strong>\u003C/h3>\u003Cp class=\"ql-align-center\" style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">Token Ops并不是独立于现有AI系统之外的新工具，而是需要落实到模型接入、请求调度和运行治理的每一个环节。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">统一接入是运营基础。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">只有将不同厂商、不同能力和不同价格的模型统一纳管，企业才能获得完整的调用数据和成本视图。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">智能调度是执行中枢。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">系统需要识别任务需求，在多个模型与供应商之间动态选择路径，实现“贵模型干贵活”，并通过负载均衡和故障切换保障业务连续性。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">计量治理形成优化闭环。\u003C/strong>\u003Cspan style=\"font-size: 12px;\">Token用量、调用成本、响应时延和成功率等数据，需要持续反馈给路由与预算策略，让模型选择和资源配置不断优化。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cstrong style=\"font-size: 12px;\">蓝耘现有能力正围绕这条路径展开：\u003C/strong>\u003Cspan style=\"font-size: 12px;\">通过\u003C/span>\u003Cstrong style=\"font-size: 12px;\">统一网关\u003C/strong>\u003Cspan style=\"font-size: 12px;\">纳管多模型资源，利用\u003C/span>\u003Cstrong style=\"font-size: 12px;\">AutoModel智能调度引擎\u003C/strong>\u003Cspan style=\"font-size: 12px;\">识别任务并选择调用路径，同时提供\u003C/span>\u003Cstrong style=\"font-size: 12px;\">Token计量、成本分析、统一鉴权、调用审计、服务监控和企业级保障\u003C/strong>\u003Cspan style=\"font-size: 12px;\">。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">这意味着，企业获得的不只是一个模型调用入口，而是一套覆盖\u003C/span>\u003Cstrong style=\"font-size: 12px;\">统一接入、智能路由、多模型调度、成本治理和稳定保障\u003C/strong>\u003Cspan style=\"font-size: 12px;\">的Token Ops运行体系，让每一次模型调用都有路径可选、有成本可查、有数据可优化。\u003C/span>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">WAIC 2026并不意味着参数竞赛结束，而是为AI竞争增加了新的评价尺度：\u003C/span>\u003Cstrong style=\"font-size: 12px;\">参数决定能力上限，Token产能体现交付效率，Token Ops决定运营价值。\u003C/strong>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cbr>\u003C/p>\u003Cp style=\"line-height: 1.5;\">\u003Cspan style=\"font-size: 12px;\">对于企业而言，真正重要的是让每一次调用都匹配合适的模型，让每一笔Token消耗都产生可衡量的结果，并将有限的算力持续转化为服务真实业务的智能产出。\u003C/span>\u003C/p>",2,"在WAIC 2026上，一个明显变化正在发生：AI产业讨论的重点，在关注「模型有多少参数」、「集群有多少GPU」的同时，开始越来越多地关注「每一份算力究竟能生产多少有效Token……",0,"2026-07-23 00:00:00",1,8,"00:00:00","2026年07月23日",1785755330490]