Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
One AI dashboard replaces your browser tabs and your Claude and GPT sub — $39.99
TL;DR: AskAnyModel AI Pro Plan puts 50+ AI models, including GPT, Claude, and Gemini, into one dashboard for $39.99 (reg. $499.00).
Comparing AI models usually means opening a bunch of browser tabs, each with its own login and interface, to see which one gives a straight answer to the same question. AskAnyModel AI Pro replaces all of that with one dashboard, and a lifetime subscription covering 50-plus AI models is on sale for $39.99 (reg. $499.00).
Sending one prompt to up to six models at once gives you a side-by-side view that shows response times and token counts next to each answer.
There are 30 models included: GPT Fast, Mistral, Llama, and DeepSeak come with unlimited use, and 500 monthly credits apply to 19 flagship models like GPT, Claude, and Gemini for tasks that require bandwidth.
A common issue with AI services that combine models is that every little thing, like making images or answering a simple question, will cost you tokens. Here, AI image generation runs on its own separate allowance, so testing a prompt across multiple image models doesn’t eat into those monthly credits.
Comparing models is easy
- Pick two to six models, including GPT, Claude, Gemini, or Grok.
- Send one prompt to every model selected at once.
- Compare the answers side by side, along with response times and token counts.
Skip the browser tabs and separate logins for every AI model. Get AskAnyModel AI Pro Plan on sale for $39.99.

AskAnyModel AI Pro Plan: Lifetime SubscriptionSee Deal
StackSocial prices subject to change.
SDUs DAISY: A Benchmark for Danish Culture
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election
EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development
Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework
How much do tech reports matter for a PhD application? [D]
The title, by tech reports I don't mean arXiv submissions, but reports of a large model, like say Kimi K3, DeepSeek, Gemini, Mistral Leanstral, etc. Is it much above, above, much below, below or equal to a first author A* paper?
[link] [comments]
Nvidia may put $10bn into Anthropic’s IPO, more than Europe’s largest AI round in full

Nvidia is in talks to invest up to $10bn in Anthropic’s initial public offering, which seeks as much as $100bn at a valuation near $2trn and would be the largest listing in history. Three days earlier it returned as an investor in Mistral’s EUR 3bn round, the largest ever raised by a European technology company […]
This story continues at The Next Web
大众拥有的第一个机器人,可能是只“鸭子”
(本文作者为 硅碳变量,钛媒体经授权发布)
文 |硅碳变量,作者 | 沈浪
最近,机器人赛道似乎跌宕起伏。
备受瞩目的宇树,上市首日股价最高涨逾6倍,总市值一度突破4400亿元;此后却连续5个交易日下跌,股价近乎腰斩,市值蒸发超过2000亿元。这一波“过山车”行情,引发了各种议论,外界对人形机器人的落地及商业化前景再次提出质疑。
而就在机器人产业迎来一盆“冷水”的时候,一只体型迷你、造型呆萌的“小鸭子”在大洋彼岸爆火,再次点燃了具身智能赛道。
它叫 Microduck,售价 399 美元——大约 2800 块人民币,是AI社区Hugging Face旗下Pollen Robotics开发推出的一款开源机器人。开订6小时,订单额破百万美元,24小时,超260万美元,高峰期每4秒卖一台,新订单排期拉到4个月。现在下单,圣诞节都未必能收到。
过去,机器人产业一直讲述“通用性”的故事,大众也在等待一款“完美人形机器人”横空出世,然而现在,在消费级市场上,Microduck(或称机器鸭)似乎给出了另一个答案。
C端比B端更早爆发?
鸭子的外形,25厘米的个头,走路一摇一摆,摔倒后又努力爬起来的样子,略显笨拙,为什么Microduck在全球能突然走红呢?
在Microduck问世前,双足机器人具身智能研发是妥妥的“高门槛赛道”。无论是波士顿动力的高端设备,还是各大厂商的全尺寸人形机器人样机,单台造价动辄数十万美金,甚至连机器狗,价格也在万元以上。而Microduck把一个能走、会动的机器人的价格拉到了千元级,这自然引得不少用户愿意尝鲜。
吸引他们的不单单是价格,而是它搭建了一套从虚拟仿真训练到真机落地部署的全流程开源强化学习管线。简单来讲就是,用户可以在仿真环境里训练“虚拟鸭子”学会一个新动作,再把训练好的行为“迁移”到实体机器人身上,自己训练自己的机器鸭。

以前机器人通常卖给了实验室、企业或高校,有了Microduck这个先例,以后机器人可能会卖给更多的开发者、极客以及普通消费者。市场也给予了Microduck热烈的反馈,仅4天,Microduck斩获10500台订单,国内某二手平台上已有代购挂出链接售卖机器鸭,价格最高达5567元,相当于原价的两倍。
Microduck走红的背后,是机器人行业正在加快向C端市场涌入。Microduck预售的前几日,上纬新材旗下启元 Q1、T1 也开启了预订,这两款产品都属于小尺寸机器人,前者定位为全球首个个人机器人,机身高度88厘米,重量15公斤,后者同样是个人机器人,但可以变形,它们都面向普通家庭。

更早一些,优必选发布的超仿生人形机器人U1系列一度在圈内引发热议,预售订单突破1.3万台。U1系列产品聚焦情感陪伴与个性化交互场景,是优必选在消费级市场的一次重要尝试。
相较于优必选U1系列和启元 Q1、T1,这款爆红的机器鸭的创造性价值不仅仅是把价格降到了普通消费者接受的范畴,还在于它为获取更高质量、多元化的真实物理交互数据提供了新思路:当海量开发者手握标准化、可二次开发的物理机器人,就能形成千万级的分布式“AI实验场景”,源源不断产出数据。
如此一来,获取数据就不再只依赖少数大型实验室或数据采集公司。
机器人赛道的主流玩家,如宇树、银河通用、小鹏、优必选等,均聚焦大尺寸人形整机研发与落地,他们想要的是创造一个通用型机器人,实现在工厂、物流等B端场景的落地。但目前为止,这条路径还未出现突破性进展,相反在C端市场,各种细分需求正在形成真实购买力,表现出意外的热情。
我们的第一个机器人,可能不是来自“宇树”们
伴随着机器人走向前沿科技的顶端、走进大众的视野,越来越多的人也产生了一种期待或者说疑问,第一台真正大规模卖给个人的机器人到底是什么样子的?它将产生于哪家机器人企业?
像Lovot、Fuzozo等机器人,尽管它们证实了陪伴、互动、情绪价值等功能,可以成为消费者买单的要素,消费者也不一定要等到机器人会干活以后才愿意购买机器人,但是产品销量已然说明这类陪伴机器人仍局限于小众范畴。更何况,现在的陪伴机器人与其说是机器人,还不如说是AI玩具。
这些AI玩具通常是把大模型装进一个毛绒玩具里,让它学着与用户聊天,产生互动,本质上其实仍停留在文本与语音的输入输出层面。
Microduck则带来了实质性的突破。作为双足机器人,它会走、会爬起、会抓取小物件,能够实现运动控制和物理交互,更接近大众理解的机器人形态。可以说,Microduck通过小尺寸本体、成熟标准件、规模采购以及标准化的硬件设计,把一台真实的机器人的门槛压到了开发者及消费者能够承受的范围。
这两年,宇树等明星级企业的出现,让机器人行业的目光主要聚焦在大尺寸的人形机器人上,可宇树上市后遭遇腰斩,人形机器人“通用性”的故事再次受到质疑,我们离拥有一个会干活的机器人似乎也更遥远了。Microduck虽然只是一只“鸭子”,可它让外界看到了机器人大规模走入日常生活的另一种可能性。
而拥有一个属于自己的、还会“成长”的个人机器人,这故事似乎也很吸引人。
不过,一个尴尬的事实是,这只火爆全球的“鸭子”,从最核心的芯片到整体的代工制造,都是Made in China,可创造了这只“鸭子”并有望开启一个开源平价机器人时代的并不是我国的任何一家机器人企业。
这意味着尽管全球机器人产业的发展离不开我国日渐成熟的机器人制造产业链,但这不代表我们就能成功定义下一代智能硬件,顺利带来电子消费市场的又一轮革新。
在机器鸭的身上,我们看到开源生态正在成为机器人行业的新变量,而Hugging Face 真正比传统机器人厂商更有优势的,正是它已经拥有大量开发者。根据Hugging Face 官方发布的 2026 年开源生态报告显示,2025 年平台用户已经增长到 1300 万,拥有超过 200 万个公开模型和 50 万个公开数据集。
Microduck仅仅是一个开始,在真正走进大众生活的硬件载体还未出现前,竞争的局势存在太多的变量。
英伟达才是最大的受益者?
就在 Microduck 开放预购的同一天,多家媒体报道英伟达已基本达成以约 129 亿美元收购 Hugging Face 的协议。如果交易完成,这将是英伟达历史上最大的收购之一。
再联想到今年年初,黄仁勋在CES主题演讲中称“下一波AI浪潮将是在物理世界中运行的AI”,这句话宣告了物理AI(Physical AI)的“ChatGPT时刻”已经到来。此时收购Hugging Face ,显然进一步释放了英伟达在具身智能领域的勃勃野心。
具体来看,从 Isaac Sim 到 GR00T 基础模型,再到 Cosmos 世界模型,英伟达已经在构建从仿真到部署的全栈机器人工具链。Hugging Face 的加入,将进一步补全开发者社区这一环,让硬件的落地门槛被进一步压低。而机器鸭的走红,意味着在英伟达的这个开源社区及背后的完整闭环中,有望诞生更多进入消费级市场的硬件,甚至是能够改变行业的爆款。
更长远地,这个开源平台聚集起大量的机器人爱好者、研究者、极客及尝鲜的消费者,这些人同样有可能成为英伟达未来的核心用户。
从布局可见,英伟达在物理AI时代的目标,是构建一个类似“安卓”之于智能手机的开放式生态系统,成为机器人及自动驾驶领域的默认开发平台。这样一来,无论机器人是继续走通用性路线,还是聚焦某一场景做专用的工业级或消费级产品,可能都离不开英伟达的芯片、模型及生态。
换句话说,别看现在机器人厂商风光无限,集体站在了聚光灯下,机器人这一风口最大的受益者或许还是英伟达。
当然,一旦某个巨头确立了主导权,对行业及其他从业者而言,往往意味着风险。就像英伟达收购Hugging Face,因为该平台同时托管 Meta、Google、Mistral 等其他公司的模型,如果归入英伟达旗下,它的中立性能否维持,是外界的疑问。
同样地,如果英伟达直接下场做机器人,便与机器人厂商形成竞争关系,拥有底层算力、自研芯片等核心优势的它,是否会形成降维打击,也是机器人厂商需要提防的。
当然这都是后话,英伟达能否如愿成为物理AI时代的“安卓”,还存在机器人何时走进生活的现实性问题。
不可否认,我们离一款“完美”的人形机器人尚且遥远。机器鸭爆火,尽管带着无数的争议,可它能让更多的人有机会亲自触摸具身智能机器人这一最前沿的技术成果,而不再是只可远观。这已然是一大进步。
更多精彩内容,关注钛媒体微信号(ID:taimeiti),或者下载钛媒体App
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
Cohere has released North Small Translate, an open-weight machine translation model from Cohere and Cohere Labs. It is a sparse Mixture-of-Experts (MoE) model with 218B total and 25B active parameters. It covers 50 languages, from Albanian to Vietnamese. On Cohere’s WMT26 evaluation, it scores 83.6 averaged across all languages. Cohere says that beats DeepL and Google Translate, plus open options like GLM 5.2 and Mistral Large 3.
Is it deployable? Yes. Call it free on Cohere’s API until rate limits, self-host it non-commercially, or license it commercially.
Back to Where the Transformer Started
Google researchers introduced the Transformer in 2017 with Attention Is All You Need. Its main results came from WMT 2014 English-to-German and English-to-French translation. 9 years later, Cohere is returning to that original problem with a dedicated model. Cohere’s launch post on X frames translation as a sovereignty issue. Organizations that cannot communicate globally cannot stay sovereign.
North Small Translate is the first translation model in Cohere’s North family. It follows Tiny Aya and Command A Translate in Cohere’s multilingual lineage. Cohere built it with RWS, whose Language Weaver scientists and language experts shaped its real-world quality.
Architecture
The model structure describes a decoder-only sparse MoE Transformer. Here are the key details:
- Experts: 128 experts, 8 activated per token, plus shared experts applied to every token.
- Router: A sigmoid over expert logits, normalized over the selected top-k.
- Attention: Sliding-window layers (window 4096, RoPE) and global layers without positional embeddings, interleaved 3:1.
- Lineage: That attention layout was first introduced in Command A.
- Context: 16K input and 16K output tokens, text only.
- Training: Post-trained specifically for translation quality.
About 11.5% of the weights are active per token. Per-token compute tracks the 25B active parameters. Memory still has to hold all 218B.
&&Benchmarks
Cohere team reports these WMT26 all-languages scores in its launch blog:
| Model | WMT26 score |
|---|---|
| North Small Translate (Agentic) | 84.36 |
| North Small Translate | 83.60 |
| Qwen 3.5 397B A17B | 81.56 |
| DeepL NextGen | 81.37 |
| Gemma 4 31B (on) | 79.46 |
| GLM 5.2 FP8 | 76.50 |
| Google Translate | 68.20 |
The Agentic variant runs a multi-pass workflow that finds and fixes its own errors. Cohere’s scoring bands treat 80 to 100 as perfect or minor errors only. One caveat matters here. These are Cohere’s own runs, with GPT-5.6-Sol as the judge. Treat them as vendor-reported until independent WMT26 results appear.
Regionally, both versions beat Gemma 4 31B (on) across Europe. On EU languages, the standard model scores 82.17 against Gemma’s 72.73. South Asia is close, at 86.16 for North against 88.04 for Gemma.
Speed, Long Documents and Cost
In Cohere’s tests, the model produced 112 output tokens per second against 81 for Gemma 4 31B. That was at low concurrency on identical hardware. At high concurrency, the figures were 39 against 30. Cohere calls this up to 1.4x higher throughput.
Long documents are a stronger point. The model scores 48.9 when translating 2 book chapters in 1 call. Google Translate scores 21.3 and Gemma 4 31B scores 19.4. Quality is measured per paragraph with xCOMET-XL.
In Cohere’s cost chart, the model scores 80.1 at $0.000676 per task, averaging 661 tokens. Gemini 3.1 Pro Preview (high) costs $0.038928 per task, about 58x more. Qwen 3.5 397B A17B costs $0.004525 and Command A+ costs $0.005158.
How to Run It
The fastest path is Cohere’s Chat V2 API. The model is free there until rate limits:
from cohere import ClientV2
co = ClientV2(api_key="<YOUR_API_KEY>")
response = co.chat(
model="north-small-translate-1-0",
messages=[{"role": "user",
"content": "Translate everything that follows into French:\n\nEnterprises need accurate translations of business-critical documents."}],
)
print(response.message.content[0].text)For self-hosting, Cohere publishes 3 checkpoints, the same ones it serves in production:
| Checkpoint | Blackwell | Hopper |
|---|---|---|
| BF16 | 4x B200 | 8x H100 |
| FP8 | 2x B200 | 4x H100 |
| NVFP4 W4A16 | 1x B200 | 2x H100 |
Key Takeaways
- Cohere’s North Small Translate is a 218B MoE with 25B active parameters.
- It scores 83.6 on WMT26 across all languages, 84.36 in agentic mode.
- All scores are vendor-reported and judged by GPT-5.6-Sol.
- The 4-bit checkpoint runs on 1x B200 or 2x H10
AI more likely to kill animals if it saves fuel or money
RDQ: Residual Distribution Quantization for Large Language Models
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
Probing for Knowledge Attribution in Large Language Models