🤖 AI 资讯

每日 05:00 更新 · 09-16 · 主站 liuch.name ↗
全部标签 →
筛选标签:办公效率 · 返回个性化推荐 · 清空筛选

从地窝子到闽宁戏剧村,西海固农妇马小琴的脱贫路

澎湃新闻
· 政策监管,办公效率,游戏,招聘HR

马上测·兼职④丨我们在律所假扮“咨询离婚诉讼”当事人,酬劳40元

澎湃新闻
· 大模型,字节跳动,办公效率,法律,招聘HR,财报,版权诉讼

马上评|一记耳光、两人受伤、三次不同判决疑问待解

澎湃新闻
· 政策监管,办公效率,教育学习,法律,招聘HR,产品更新,版权诉讼

释新闻|三地突然闹“脱英”,英国真的面临“解体危机”?

澎湃新闻
· 政策监管,办公效率,医疗健康,教育学习,法律,政务,招聘HR,模型发布,合作,版权诉讼

分析|怎么看8月消费、投资增速放缓?

澎湃新闻
· 算力芯片,AI应用,融资,政策监管,办公效率,医疗健康,金融,政务,工业制造,零售电商,营销广告,AI for Science,招聘HR,基础设施,产品更新,财报
AI 资讯

Trying to Make a Loop Auto-Vectorize

Lobsters

Comments

· 具身智能,办公效率,招聘HR
AI 资讯

Learning to solve hard problems in RL for LLMs by never giving up

Hacker NewsComments
· 大模型,AI应用,开源,OpenAI,阿里巴巴,DeepSeek,Agent智能体,推理思考,搜索RAG,办公效率,强化学习,微调蒸馏,模型评测,提示工程,招聘HR,榜单评测,论文

We got admin access to Baseten's production GitHub

Hacker NewsComments
2026-09-01 · AI应用,开源,Meta,代码生成,Agent智能体,搜索RAG,办公效率,招聘HR,榜单评测
AI 资讯

Universal Feature Selection with Noisy Observations and Weak Symmetry Conditions

arXiv cs.LGarXiv:2605.09396v2 Announce Type: replace-cross Abstract: This paper relaxes the restrictive symmetry conditions adopted in [4], [5] and extends their universal feature selection framework to accommodate noisy observations as well as attribute structures that may exhibit directional preferences. We introduce the notion of weak spherical symmetry, quantified by second-moment distances, which allows controlled deviations from rotational invariance. Under this relaxed condition, we develop a universal feature selection framework based on the singular value decomposition of the canonical dependence matrix computed from noisy data. Our main result shows that the selected features achieve asymptotically optimal error exponents up to a residual term that depends on the symmetry deviation $\delta$ and the noise levels $\eta_1, \eta_2$. When $\delta, \eta_1, \eta_2$ are relatively small, our result recovers that of [5], thereby demonstrating that exact spherical symmetry is unnecessary. Overall, our findings highlight the robustness of the selection framework against second-moment deviations and observation noise, thereby broadening its applicability across diverse inference tasks and providing a theoretically grounded tool for universal feature selection in practical scenarios.
2026-09-16 04:00:00 · 办公效率,扩散模型,论文
AI 资讯

Structure-Aware Masking for Protein Representation Learning

arXiv cs.LGarXiv:2605.16581v2 Announce Type: replace Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual residues at a fixed rate (e.g., 15%). This practice implicitly assumes that all sequence positions contribute equally to representation learning. In downstream fitness prediction tasks, however, protein sequences are governed by three-dimensional structural dependencies and long-range residue contacts that induce strong nonlocal couplings between residues. We introduce Bucket Masking, a structure-aware masking strategy that selects groups of residues based on their proximity in three-dimensional space, preferentially masking structurally coupled regions during training. By conditioning the masking distribution on residue contacts, Bucket Masking shifts the learning objective toward modeling long-range interactions that are critical for protein function. Across four downstream protein fitness prediction tasks, Bucket Masking enables up to a 14% improvement over standard random masking, excelling at predicting higher-order mutational interactions. Through controlled ablations, we show that these improvements arise from mask placement rather than span size, establishing masking as a positional inductive bias.
2026-09-16 04:00:00 · 办公效率,扩散模型,招聘HR,论文
AI 资讯

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

arXiv cs.LGarXiv:2609.16606v1 Announce Type: new Abstract: Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure. TSKs parameterize the weights on each multivariable component of the target function by factors for each input. We propose learning these factors directly from function evaluations by selecting the reproducing kernel Hilbert space (RKHS) in which the target function has minimum norm. Under suitable conditions, we show that this norm-minimization problem admits a unique solution, and we establish consistency of a finite-data formulation based on minimum-norm interpolation. The learned TSK factors characterize the participation of individual inputs across interactions and main effects, providing a kernel-dependent notion of input sensitivity related to total Sobol indices. Numerical experiments demonstrate that adapting the kernel to learned multivariable structure can substantially improve approximation accuracy over a standard product kernel.
2026-09-16 04:00:00 · 算力芯片,Google,办公效率,扩散模型,端侧AI,论文
AI 资讯

What Does Layer-Importance Reveal About Transformers and State-Space Models?

arXiv cs.LGarXiv:2609.16537v1 Announce Type: new Abstract: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. \emph{Necessity} captures how much the pretrained model depends on a layer's existing contribution, measured by the loss increase from bypassing it. \emph{Plasticity} captures where the model absorbs new information during fine-tuning, measured by the magnitude of task-specific weight updates. Our analysis reveals that the two families behave fundamentally differently: in every evaluated residual transformer up to $14$B parameters, Necessity and Plasticity anti-align across depth, whereas in the evaluated Mamba-style SSMs they point to overlapping regions. The sign of this alignment also predicts downstream adaptation behavior. In the evaluated transformers, concentrating updates in the most plastic layers increases catastrophic forgetting, while this tier-dependent effect disappears in the evaluated Mamba-style SSMs.
2026-09-16 04:00:00 · 办公效率,Transformer,微调蒸馏,预训练,模型安全对齐,招聘HR,论文
AI 资讯

How Good Are Time-Series Foundation Models for Pedestrian Crowd Count Forecasting? A Cross-Dataset Comparative Study

arXiv cs.LGarXiv:2609.16415v1 Announce Type: new Abstract: Pedestrian-count forecasting supports pedestrian-oriented Intelligent Transportation Systems (ITS), including crowd monitoring, pedestrian-traffic staffing and routing, and proactive risk mitigation during surges. Recent time-series foundation models (FMs) report strong zero-shot accuracy on heterogeneous forecasting benchmarks, but it remains unclear whether these gains transfer reliably to pedestrian sensing deployments. We benchmark seven univariate forecasting approaches spanning four paradigms: Seasonal Naive, gradient-boosted trees (LightGBM, CatBoost), deep learning models (N-HiTS, PatchTST), and two pretrained FMs (TimesFM, Chronos-2). Experiments cover two complementary regimes: (i) a five-day special event dataset SAIL2025 at 3-minute resolution with limited in-domain history; and (ii) Melbourne pedestrian sensors as a multi-year hourly dataset (2010--2017) with strong seasonality. We compare the MAE and RMSE results per sensor across datasets and multiple forecast horizons. Results show three consistent findings. First, with limited historical data, Seasonal Naive remains a strong baseline for long-horizon forecasting on high-volume sensors, while trained models can degrade when the next day differs substantially from prior days. Second, boosted trees can be competitive on lower-volume sensors but exhibit higher sensitivity on high-volume sensors under event-driven shift. Third, FMs excel in the seasonal and data-rich regime under long-context configuration. The findings highlight the importance of choosing pedestrian forecasting models based on both the underlying data conditions and the forecasting horizon.
2026-09-16 04:00:00 · 办公效率,扩散模型,强化学习,预训练,模型评测,招聘HR,论文
AI 资讯

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

arXiv cs.CVarXiv:2609.16011v1 Announce Type: cross Abstract: Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produce only linguistic outputs, leaving a critical gap in embodied response generation. We identify and address a failure of emotion conditioning: like other conditional generators that under-use weak conditioning signals, a flow-matching model given both a rich audio embedding and a discrete emotion label suppresses the emotion, generating near-identical motion regardless of the specified emotion. We present EMODY Flow, a lightweight (around 35M parameters) flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions. A training-time auxiliary emotion classifier restores emotion sensitivity by forcing generated motion to be emotion-identifiable. EMODY Flow sets a new state of the art on BEAT2 gesture quality, with FGD 0.302, Beat Correlation 0.853, and Diversity 24.62 - improving over the best prior results by 26%, 5%, and 62% respectively - and transfers to zero-shot facial animation on TFHP without domain-specific fine-tuning. Beyond these quantitative gains, the classifier yields clearly emotion-separated motion, which we demonstrate qualitatively through a multidimensional-scaling analysis of the generated gestures.
2026-09-16 04:00:00 · 大模型,算力芯片,AI应用,具身智能,Google,阿里巴巴,多模态,Agent智能体,办公效率,扩散模型,微调蒸馏,向量数据库,招聘HR,论文
AI 资讯

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

arXiv cs.CVarXiv:2609.16233v1 Announce Type: new Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.
2026-09-16 04:00:00 · 算力芯片,推理思考,办公效率,模型评测,招聘HR,论文
AI 资讯

Generating Individual Travel Diaries Using Large Language Models Informed by Census and Land-Use Data

arXiv cs.CLarXiv:2509.09710v3 Announce Type: replace Abstract: This study introduces a Large Language Model (LLM) scheme for generating key attributes of travel diaries in agent-based transportation models, including purpose, mode and distance, to assess the underlying viability of LLMs for activity generation tasks. While traditional approaches rely on large quantities of proprietary household travel surveys, our method generates personas stochastically from open-source American Community Survey (ACS) and Smart Location Database (SLD) data, then synthesizes diaries through direct prompting. Our study features a novel one-to-cohort realism score: a composite of four metrics (Trip Count Score, Interval Score, Purpose Score, and Mode Score) validated against the Connecticut Statewide Transportation Study (CSTS) diaries, matched across demographic variables. Our validation utilizes Jensen-Shannon Divergence to measure distributional similarities between generated and real diaries. When compared to diaries generated with classical methods (Negative Binomial for trip generation; Multinomial Logit for mode/purpose) calibrated on the validation set, LLM generated diaries achieve comparable overall realism (LLM mean: 0.692 vs. 0.628). The LLM excels in determining trip purpose, and its trip mode predictions demonstrate greater consistency (a narrower Realism Score distribution). Meanwhile, classical models lead to better numerical estimates of trip count and activity duration. Aggregate validation confirms the LLM's statistical representativeness (LLM mean: 0.779 vs. 0.706), demonstrating LLM's zero-shot viability and establishing a quantifiable metric of diary realism for future synthetic diary evaluation systems.
2026-09-16 04:00:00 · 大模型,AI应用,Agent智能体,办公效率,扩散模型,提示工程,招聘HR,论文
AI 资讯

The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis

arXiv cs.CLarXiv:2609.16076v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework that leverages executable code as a rigorous medium for reasoning data synthesis. MIMIC fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation. Crucially, these explicit intermediate execution states naturally form a Code-Instrumented Reward (CIR), providing dense, high-fidelity process supervision for reinforcement learning without external reward models. Extensive evaluations reveal that models trained via SFT and GRPO on our synthesized dataset achieve substantial, consistent gains. Our method significantly elevates accuracy across general reasoning, complex mathematical benchmarks, and fine-grained deterministic tasks, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabilities of LLMs. Our code and data are available at https://github.com/zjy1298/MIMIC.
2026-09-16 04:00:00 · 大模型,AI应用,开源,推理思考,搜索RAG,办公效率,强化学习,模型评测,招聘HR,论文

‘Now We Can Know Everything and Do Anything,’ Jensen Huang Says at Dreamforce

NVIDIA Blog

Know everything. Do anything.

That was the message NVIDIA founder and CEO Jensen Huang brought to Salesforce Dreamforce Tuesday, joining CEO Marc Benioff onstage in an appearance that coincided with the announcement of Koa — Salesforce’s first CRM reasoning model, built on NVIDIA Nemotron 3 Super.

Huang didn’t just take the stage. He walked into the crowd — threading through the audience at Moscone Center, mic in hand, with Benioff alongside. “We’re all following you, bro,” Benioff told Huang at one point, drawing applause. 

“With electricity, we could power everything,” Huang said, tracing the arc of industrial revolutions. “With the internet, you can find anything. And now, with artificial intelligence as an infrastructure layer across the planet, we can know everything and do anything.”

That arc led directly to the announcement. Salesforce’s first reasoning model was built by post-training NVIDIA Nemotron 3 Super on a proprietary synthetic dataset drawn from nearly three decades of enterprise CRM deployments.

AI Safety

When Benioff asked Huang about AI safety, Huang was direct. 

“Safety is paramount in a lot of ways. It’s job one,” Huang said. “However, safety is an engineering problem. We’re developing computing systems after all … If you build a product or a service and you’re not confident in its functionality, capability or safety, then don’t release it.”

He reiterated that innovation and leadership should coincide with safety and security. “Innovation, speed and safe products — it’s a false choice,” Huang said. “You could definitely have both at the same time. So run as fast as you can. But if you feel at any given point in time the company’s out of control, or the product’s not going to be safe, take a pause and make sure you get it right.”

Throughout his talk, Huang offered a positive vision for the future of AI. He noted he was among the first to push back on the notion that AI would end the software industry. He also pushed back on the idea that AI is a threat to jobs. 

“As a result of our ambition, and with the productivity boost we get from AI, the sky’s the limit,” he added. “Engage AI. Don’t get left behind.”

“Every company, every enterprise, every country would become an AI company,” Huang said. “It’s going to be an agentic enterprise.”

“The sky is especially the limit for our Trailblazers,” Benioff said, referring to the company’s customer and developer community. “I think this technology, the way it empowers people … the technology can partner with them to do this in incredible new ways.”

Because NVIDIA Nemotron is open, Salesforce was able to fine-tune and run the model entirely within its own infrastructure. Salesforce controls the weights, the model runs on Salesforce’s systems and no customer data is used during training or inference.

“Not a single byte of customer data was used,” said Rohan Kumar, Salesforce’s president of platform and engineering, who introduced Koa before Huang took the stage.

Open models went from 30% at the beginning of last year to now some 70%, Huang explained.

“People are both adopting closed models at exponential rates, but also building their own custom AIs — because every single software company is an AI company, and every single enterprise is going to be an agentic and AI enterprise.”

NVIDIA is one of them — running on Salesforce across sales, service, marketing and operations, and currently piloting Agentforce for customer-support workflows.

Koa was built using supervised fine-tuning and reinforcement learning with NVIDIA NeMo RL, NeMo Gym and NeMo AutoModel. 

The training corpus covers synthetic enterprise scenarios across 14+ industries — manufacturing, financial services, healthcare and travel. 

In Salesforce’s CRM Bench, a model benchmark that includes a suite of real-world tasks like updating an opportunity, routing a case, or scheduling a follow-up, Koa already matches or exceeds leading model performance on CRM actions with 3x fewer errors.

In Use Now

Koa is already running inside Salesforce, powering an employee agent in Slack, and is moving into customer pilots in October as a customer-selectable model in Agentforce, starting with Formula 1, UChicago Medicine, Baxter Credit Union, 1-800Accountant, Engine and Xero. General availability is expected winter 2026 in U.S. regions.

2026-09-15 22:24:34 · 算力芯片,AI应用,NVIDIA,Agent智能体,推理思考,办公效率,扩散模型,强化学习,微调蒸馏,模型评测,招聘HR

12 overlooked PC building tools I wish I’d started using years ago

PCWorld

I’ve been building PCs for over 20 years, and though I’ve gotten pretty good at it, a big part of that is due to the tools I use.

Everyone knows you need a decent Phillips screwdriver and enough space to build comfortably, but over the years and many builds, I’ve acquired a collection of extras that make the whole process even smoother.

From improved lighting to magnetic trays for all your screws to a blob of sticky tack that’s surprisingly useful, here are my picks for the most overlooked PC building tools that should be in your kit.

LED flashlight gloves

Hinshark LED flashlight gloves

Hinshark

If you’ve built even a single PC, you’ll know that lighting is a real problem. Black PC cases with lots of nooks and crannies make it hard to see what you’re doing—even with good overhead lighting. Angle lamps are great, but I’m a big fan of hands-free lighting that can peer into any bit of darkness as needed, ideally without stopping what I’m doing.

Some people use headlamps, but I find these LED flashlight gloves work a lot better. Yes, you do need to point your fingers exactly where you want the light to go, but that’s fine since you’re already deep in the case while working on it. It lets you aim the light at all those weird angles that could normally stump you, and it’s excellent for locating screws that fall behind the motherboard tray or PSU shroud.

Get the Hinshark LED Flashlight Gloves at Amazon

Needle-nose pliers

Workpro Needle-Nose Pliers

Workpro

My hands aren’t the biggest, but even my fingers can’t get into some of the tinier confines of modern desktop PCs and laptops. What do you do when you’ve lost screws deep within a case? One option is to lift the whole thing up and shake it upside-down like it owes you money. My preference is to use needle-nose pliers—or in a “pinch,” a hemostat—for reaching into those places where your fingers simply can’t.

Get the Workpro Needle-Nose Pliers at Amazon

Extra-long magnetic screwdriver

Extra Long-Neck Magnetic Screwdriver

Kyuionty

You already have a screwdriver, but do you have a long-neck model with magnetic tip? A standard screwdriver is fine most of the time, but this is a great supplementary option that can really save your bacon in certain cases. The long neck makes it easier to screw stuff in without having to reach all the way into the case, and the magnetic tip makes screw recovery from hard-to-reach cracks easier.

Get this extra-long magnetic screwdriver at Amazon

Screw assortment kit

Kernmax Computer Screw Assortment Kit
2026-09-15 13:00:00 · AI应用,具身智能,搜索RAG,办公效率,扩散模型,招聘HR,收购并购,榜单评测
继续滚动加载更多…