🤖 AI 资讯

每日 05:00 更新 · 09-16 · 主站 liuch.name ↗
全部标签 →
筛选标签:论文 · 返回个性化推荐 · 清空筛选

RSI is not happening [R]

Reddit r/MachineLearning
RSI is not happening [R]

A new paper (I'm not a coauthor BTW -- I just found it interesting) argues, basically, that RSI1 is not on the horizon2, because current (at the time the study was done) agents cannot do open-ended ML research.

Specifically, they took some accepted, but unpublished papers from NeurIPS, and tried to get the agents to do the same work, which was then graded by the original authors. And the agents (Codex/GPT-5.6 Sol and OpenClaw/Opus 4.8) could not do it.

And since they cannot do open-ended ML research, they cannot recursively self-improve -- this is their argument.3

Link: https://arxiv.org/abs/2607.27191

I think I've regretted the last 10 or so times I posted any kind of "research" in this subreddit -- either people downvote it, or it gets upvoted, but there is zero meaningful discussion. This might be the last time I'm trying this.4

Footnotes:

  1. RSI = Recursive self-improvement, a.k.a. superintelligence explosion. The concept was invented by I.J. Good in 1965. It does not mean "anything that speeds up AI research". Compilers speed it up! RSI means, basically, a nuclear chain-reaction, but for AI. The paper talks about "explosive AI progress" in the very first sentence of the abstract, and mentions "RSI" in the text.
  2. Some people have objected to my use of "X is on the horizon". I consider it synonymous to "people forecast X", and the authors use the word "forecast". "Not on the horizon" does not mean "can never happen".
  3. Quote: "This design also allows us to test a mechanism that informs many forecasts of recursive self-improvement: AI agents accelerate AI research because researchers delegate entire projects to agents and judge whether the returned results advance their work. Our evaluation closely matches this model, since authors handed an agent their own research question and closely evaluated the resulting output."
  4. 3 years ago, many of you upvoted a bunch of very uninformed comments that accused some researchers of misconduct, until I explained that this stemmed from misunderstanding how training works, in practice: https://www.reddit.com/r/MachineLearning/comments/18bdcu7/r_sequential_modeling_enables_scalable_learning/kc60k7e/?context=3 Today, one of the top comments is "I read the abstract (...) Nowhere, absolutely nowhere, do they make the claim ...". It's completely absurd. (Also, the commenter doesn't understand what "RSI" means.) The hivemind is very disappointing.

https://preview.redd.it/kprgucsxaoph1.png?width=796&format=png&auto=webp&s=60fdf26d150d9588e90b3d08e6a1b7fd84192ba5

submitted by /u/we_are_mammals
[link] [comments]
2026-09-14 18:03:41 · 大模型,算力芯片,AI应用,OpenAI,Google,代码生成,Agent智能体,扩散模型,强化学习,端侧AI,招聘HR,榜单评测,论文
AI 资讯

Learning to solve hard problems in RL for LLMs by never giving up

Hacker NewsComments
· 大模型,AI应用,开源,OpenAI,阿里巴巴,DeepSeek,Agent智能体,推理思考,搜索RAG,办公效率,强化学习,微调蒸馏,模型评测,提示工程,招聘HR,榜单评测,论文
AI 资讯

Information-Theoretic Bounds for Sparse Covariance Estimation in the Vertical-Split Distributed Model

arXiv stat.MLarXiv:2606.07124v2 Announce Type: replace-cross Abstract: We study the minimax estimation error for distributed covariance matrix estimation in the vertical-split (feature-split) setting, where two agents each observe different coordinates of~$m$ i.i.d.\ sub-Gaussian samples and communicate a limited number of bits to a central server. While \cite{rahmani2025fundamental} established nearly tight bounds for dense (unstructured) cross-covariance matrices, we investigate whether imposing elementwise $s$-sparsity on the cross-covariance $C_{21}$ can reduce the required communication and sample complexity. In contrast to the horizontal-split setting, where \cite{braverman2016communication} showed that sparsity does \emph{not} reduce communication cost for mean estimation, we prove that sparsity \emph{does} help for cross-covariance estimation in the vertical split. Specifically, for sufficiently large $d_1d_2/s'$ and $0<\varepsilon<\sigma^2\sqrt{s'}/32$, any scheme achieving expected Frobenius distortion at most $\varepsilon$ must satisfy $B_k = \Omega(\sigma^4 d_k\, s' \log(d_1 d_2/s')/\varepsilon^2)$ and $m = \Omega(\sigma^4\, s' \log(d_1 d_2/s')/\varepsilon^2)$ for cross-covariance estimation, where $s' = s \wedge d_{\min}$. For the $1$-sparse case, our achievable scheme reduces the $d_1d_2$ factor in the dense communication rate to $\log(d_1d_2)$, up to polylogarithmic factors, for the cross-covariance communication component in the matching regime. Our lower bounds are established via Fano's method with an explicit sparse packing using a Varshamov--Gilbert-type argument for signed partial permutation matrices combined with the Conditional Strong Data Processing Inequality of \cite{rahmani2025fundamental}. We show that the communication lower bound is tight up to polylogarithmic factors under the conditions of Remark~\ref{rem:achievmatch}, using an achievable scheme based on covering-net quantization and entry-wise hard thresholding.
2026-09-16 04:00:00 · 大模型,AI应用,Agent智能体,扩散模型,微调蒸馏,招聘HR,论文
AI 资讯

Statistical Inference for Score Decompositions

arXiv stat.MLarXiv:2603.04275v2 Announce Type: replace-cross Abstract: We introduce inference methods for score decompositions, which partition scoring functions for predictive assessment into three interpretable components: miscalibration, discrimination, and uncertainty. Our estimation and inference relies on a linear recalibration of the forecasts and is applicable to general point forecasts such as means and quantiles due to its validity for non-smooth scoring functions. This approach ensures non-negative decomposition terms in finite samples, enables asymptotic inference under model misspecification, and establishes a direct connection to the classical Mincer-Zarnowitz regression. The resulting inference framework facilitates novel tests for equal linearized forecast calibration or discrimination, which yield three key advantages. They enhance the information content of predictive ability tests by decomposing scores, can improve detection power in scenarios where predictive differences are attributable to specific score components, and formally connect scoring-function-based evaluation to traditional calibration tests, such as financial backtests. Applications demonstrate the method's utility. We find that for survey inflation forecasts, discrimination abilities can differ significantly even when overall predictive ability does not. In an application to financial risk models, our tests provide deeper insights into the calibration and information content of volatility and Value-at-Risk forecasts. By disentangling forecast accuracy from backtest performance, the method exposes critical shortcomings in current banking regulation.
2026-09-16 04:00:00 · 扩散模型,招聘HR,论文
AI 资讯

A proximal augmented Lagrangian method for nonconvex optimization with equality and inequality constraints

arXiv stat.MLarXiv:2509.02894v2 Announce Type: replace-cross Abstract: We propose an inexact proximal augmented Lagrangian method (P-ALM) for nonconvex structured optimization problems. The proposed method features an easily implementable rule not only for updating the penalty parameters, but also for adaptively tuning the proximal term. It allows the penalty parameter to grow rapidly in the early stages to speed up progress, while ameliorating the issue of ill-conditioning in later iterations, a well-known drawback of the traditional approach of linearly increasing the penalty parameters. A key element in our analysis lies in the observation that the augmented Lagrangian can be controlled effectively along the iterates, provided an initial feasible point is available. Our analysis, while simple, provides a new theoretical perspective about P-ALM and, as a by-product, results in similar convergence properties for its non-proximal variant, the classical augmented Lagrangian method (ALM). Numerical experiments, including convex and nonconvex problem instances, demonstrate the effectiveness of our approach.
2026-09-16 04:00:00 · 扩散模型,论文,开发者生态
AI 资讯

Kinetic Interacting Particle Langevin Monte Carlo

arXiv stat.MLarXiv:2407.05790v4 Announce Type: replace-cross Abstract: This paper introduces and analyses interacting underdamped Langevin algorithms, termed Kinetic Interacting Particle Langevin Monte Carlo (KIPLMC) methods, for statistical inference in latent variable models. We propose a diffusion process that evolves jointly in the space of parameters and latent variables and show that the stationary distribution of this diffusion concentrates around the maximum marginal likelihood estimate of the parameters. We then provide two explicit discretisations of this diffusion as practical algorithms to estimate parameters of statistical models. For each algorithm, we obtain nonasymptotic rates of convergence in Wasserstein-2 distance for the case where the joint log-likelihood is strongly concave with respect to latent variables and parameters. We achieve accelerated convergence rates clearly demonstrating improvement in dimension dependence. To demonstrate the utility of the introduced methodology, we provide numerical experiments that illustrate the effectiveness of the proposed diffusion for statistical inference. Our setting covers a broad number of applications, including unsupervised learning, statistical inference, and inverse problems.
2026-09-16 04:00:00 · 扩散模型,论文
AI 资讯

Covariate Selection for Doubly Robust Double/debiased Machine Learning Estimators for Causal Inference

arXiv stat.MLarXiv:2609.17238v1 Announce Type: cross Abstract: High-dimensional data create challenges for causal effect estimation because identifying the covariates needed for correct model specification becomes increasingly difficult. Double/debiased machine learning (DML) facilitates the use of machine learning (ML) for causal inference by mitigating regularization and overfitting bias, but comparatively less attention has been given to covariate selection in relation to the double robustness (DR) property possessed by some DML estimators. In particular, ML-based covariate selection may result in differential covariate selection or in misspecification of both models, thereby limiting the practical utility of the DR property. To address these issues, we propose using the union of the covariates selected by the propensity score (PS) and outcome ML models to re-estimate both models. Simulation results show that using the union consistently reduces more confounding bias than using separate selected covariate sets. The results also show that ML-based estimation does not uniformly outperform conventional DR estimation, even under conditions favorable to the Lasso, and that post-Lasso reduces more confounding bias than standard Lasso. These findings demonstrate that successful use of ML for causal inference depends not only on the ML algorithm but also on how the information obtained through covariate selection is incorporated into causal effect estimation.
2026-09-16 04:00:00 · Transformer,扩散模型,招聘HR,论文
AI 资讯

Learning Choice Model Trees for Feature-Based Multi-Product Pricing: Exact Optimization and Field Evidence

arXiv stat.MLarXiv:2609.16952v1 Announce Type: cross Abstract: Feature-based multi-product pricing uses customer characteristics to identify demand heterogeneity and tailor prices across products. Choice model trees segment customers through interpretable feature rules and fit a demand model within each leaf. Existing methods typically construct these trees greedily, selecting one myopic split at a time. We develop optimal choice model trees with multinomial logit leaves (OCMT-MNL), jointly optimizing the tree and leaf models within a prescribed depth. Our exact dynamic program derives closed-form Fenchel lower bounds during constrained Newton iterations and propagates them across nested and disjoint customer subsets, avoiding new fits and resuming unfinished fits without repeating completed work. In synthetic experiments, it reduces exact leaf fits by 99.98% and leaf evaluations by 86.13%, achieving up to 7.15-fold speedups over unpruned dynamic programming. One-dimensional lookup tables translate offline estimation into real-time pricing, with a revenue-loss bound quadratic in grid spacing under the fitted model. Compared with greedy trees, OCMT-MNL achieves lower revenue loss with fewer leaves on synthetic data and better predictive fit on real data. In a 23-week randomized experiment on ancillary seat pricing across 48 airline markets and 190,220 passengers, OCMT-MNL increases seat revenue per passenger by a statistically significant 11.3% over static pricing.
2026-09-16 04:00:00 · 招聘HR,论文
AI 资讯

Characterizing Heterogeneous Rates in Finite Mixture Estimation via Partial Optimal Transport

arXiv stat.MLarXiv:2609.16622v1 Announce Type: cross Abstract: Parameter estimation in finite mixture models can exhibit highly heterogeneous convergence behavior: locally isolated components may be estimated substantially faster than groups of competing components. Existing analyses based on Wasserstein distances typically characterize only the worst-case rate and therefore do not fully capture this local heterogeneity. In this paper, we introduce a Voronoi-based partial optimal transport (VPOT) framework for obtaining refined local and global convergence guarantees for the maximum likelihood estimator of the mixing measure. The key geometric idea is to localize the comparison of two mixing measures to extended Voronoi neighborhoods and use partial optimal transport to accommodate the unequal masses of their local restrictions. Within each neighborhood, the first-order POT discrepancy is raised to a power determined by the number of locally competing atoms, allowing the resulting loss to adapt to the local degree of singularity. Under suitable regularity and strong identifiability conditions, we establish uniform local and global upper bounds for a maximum likelihood estimator under the VPOT loss. These bounds reveal a configuration-dependent form of parameter estimation: less singular local configurations admit faster convergence, whereas the most singular configuration recovers the classical worst-case behavior characterized by Wasserstein-based analyses. We further establish a minimax lower bound showing that the convergence rate for estimating the mixing measure under the VPOT loss is optimal. Our results hold in arbitrary fixed dimension without requiring mixing proportions to be uniformly bounded away from zero or prior knowledge of the true number of mixture components. Overall, VPOT provides a configuration-adaptive framework for capturing heterogeneous parameter-estimation behavior in finite mixture models.
2026-09-16 04:00:00 · 大模型,扩散模型,论文
AI 资讯

Supervising the Chain Ladder

arXiv stat.MLarXiv:2609.16552v1 Announce Type: cross Abstract: The chain ladder's volume-weighted pattern minimises an explicit loss function, yet is rarely booked as such. Practitioners adjust the pattern and record the final adjusted ratios. This paper treats the chain ladder's pattern selection as a supervised-learning problem. Judgement on pattern adjustments becomes a framework of defined penalties and hyperparameters on the chain ladder's loss function, treated here as an objective function in machine learning. Data weights are generalised with a decay and a power parameter for recency and volume weighting. Benchmark shaping and smoothness enter through a reference penalty and Whittaker-Henderson smoothing. The assembled objective is strictly convex and minimised by a single linear system. Each hyperparameter becomes an interpretable adjustment in its own right, declarable by judgement and categorised as an experience or a prospective adjustment. Experience adjustments can be set more objectively by a proposed training loop and a reserve validation score on held-out calendar diagonals. Further hyperparameter-based adjustments are written as almost-everywhere differentiable penalties that re-time or reshape the pattern. A worked example carries one real Schedule P triangle through an incurred and then a paid training stage, demonstrating the workflow.
2026-09-16 04:00:00 · 模型评测,招聘HR,论文,开发者生态
AI 资讯

Causal Path Analysis from Perturbational and Population-Scale Single-Cell Data with Multiscale Confounding and Measurement Error

arXiv stat.MLarXiv:2609.16510v1 Announce Type: cross Abstract: Single-cell perturbation experiments provide causal information on gene regulation, whereas population-scale single-cell studies characterize gene expression and phenotypes in human populations. We develop a framework that integrates these complementary data sources for causal path analysis. Rather than assuming that a perturbational gene network transfers directly to the target population, we use externally learned ancestral relationships to constrain the network topology and re-estimate its direct edges and effects from population data. To address latent heterogeneity and measurement error in multiscale single-cell measurements, we develop a surrogate-variable procedure operating at both the cell and subject levels, combined with errors-in-variables correction for network and outcome regressions. We establish theoretical guarantees for confounder recovery and high-dimensional estimation of network and gene-outcome effects. Simulations demonstrate the importance of jointly correcting confounding and measurement error. An application to acute myeloid leukemia identifies distinct regulatory pathways linking transcriptional regulators to blast count.
2026-09-16 04:00:00 · 榜单评测,论文
AI 资讯

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold

arXiv cs.LGarXiv:2608.18402v2 Announce Type: replace-cross Abstract: We study finite-sample linear regression in the presence of varied and unknown label noise, focusing on the heteroskedastic and adaptive linear regression models. Heteroskedastic linear regression models settings where the labels are of varying quality. We receive $n$ pairs $(X_i,Y_i)$ with labels $Y_i=X_i^\top\beta+\varepsilon_i$, where $\varepsilon_i\sim N(0,\sigma_i^2)$ and the variances are unknown to the estimator. One natural measurement of the difficulty of this problem is the number of samples $m$ for which $\sigma_i^2\le1$ (larger $m$ is easier). We obtain a polynomial-time estimator with rate $\tilde{O}((nd^3/m^4)^{1/6})$ when $m\gg d^{3/4}n^{1/4}$, as well as nearly-matching lower bounds. For $d=O(1)$, our estimator achieves error $o(1)$ when $m\gg n^{1/4}$, whereas $L_1$ regression and other traditional approaches require $m\gg n^{1/2}$. In adaptive linear regression, the errors are drawn i.i.d. from an unknown distribution $p$, and our goal is to design a generic estimator that performs nearly as well as the best custom estimator that knows $p$. We introduce a (computationally inefficient) adaptive estimator that, so long as $p$ is a mixture of $k$ symmetric log-concave densities, achieves error comparable with the optimal estimator that knows $p$ and has $\tilde\Theta(n/k)$ samples. For $k=1$, we show that $L_q$ regression (with data-dependent $q$) gives a polynomial-time estimator. Finally, to study the computational limits of both problems, we introduce the planted linear regression problem, where $X_i\sim N(0,I_d)$, $m$ unknown samples are noiseless, and the rest have error $\varepsilon_i\sim N(0,1)$. We conjecture that recovering $\beta$ up to error $\ll\sqrt{d/n}$ (or exactly) may have an information-computation gap between $m=d+1$ and $m\sim d^{3/4}n^{1/4}$, as is suggested by our near-matching polynomial-time estimator and statistical query (SQ) lower bound.
2026-09-16 04:00:00 · 扩散模型,招聘HR,榜单评测,论文
AI 资讯

Protecting patient privacy in clinical foundation models: Technical and legal perspectives

arXiv cs.LGarXiv:2608.07705v2 Announce Type: replace-cross Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health planning. As deployment expands, privacy risk arises from model-mediated leakage, yet its prevalence and severity remain poorly quantified. Models can disclose sensitive training artifacts, enabling patient re-identification in ways not captured by data-handling controls alone. As a result, existing frameworks, including HIPAA and GDPR, offer limited protection against assessing and addressing. We propose a practical framework for assessing privacy risk in clinical foundation models, illustrate realistic leakage scenarios across deployment settings, map them to legal regimes, and outline complementary technical and legal mitigations. Our analysis provides a context-aware risk assessment grounded in realistic usage to preserve the value of medical foundation models while rigorously safeguarding patient privacy.
2026-09-16 04:00:00 · 强化学习,论文
AI 资讯

Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

arXiv cs.LGarXiv:2607.19837v2 Announce Type: replace-cross Abstract: Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment. We formalize agent reconnaissance by modeling the process and identifying the knowledge assets it seeks to extract: what they are, how they are used, and which agent weaknesses they exploit to give adversaries leverage in indirect prompt injection attacks. We instantiate these insights in Know Your Agent (KYA), a framework that automates black-box, reconnaissance-driven pentesting by probing agents, building target profiles, and using those profiles to craft stronger attacks. We evaluate KYA on agent-security benchmarks and a real-world coding agent, and release KYA, its benchmarks, and baseline implementations for reproducibility.
2026-09-16 04:00:00 · AI应用,Agent智能体,搜索RAG,扩散模型,模型评测,提示工程,论文
AI 资讯

Missing Data Imputation under Manifold Hypothesis

arXiv cs.LGarXiv:2607.03641v3 Announce Type: replace-cross Abstract: The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold. Recent advances in mixture variational autoencoders (VAEs) provide a powerful tool for extracting such underlying structure in a faithful manner. The resulting geometric structure naturally introduces local and global relationships among variables, thereby providing a systematic way of imputing missing data. We propose a model-based imputation method that enables sampling from \( p(\bm{x}_{\mathrm{mis}} \mid \bm{x}_{\mathrm{obs}}) \) via a sampling-importance-resampling (SIR) procedure, which can be further augmented with a joint diffusion model in the latent space. Our method imputes missing data while respecting the underlying geometry, achieves competitive performance compared to state-of-the-art procedures, quantifies uncertainty in the imputations, and is model-based, thereby enabling on-the-fly imputation without rerunning the entire procedure.
2026-09-16 04:00:00 · 扩散模型,招聘HR,论文
AI 资讯

Shielded Analysis: Certification and Characterization of Defensibility in Systems under Adversarial Interaction

arXiv cs.LGarXiv:2606.13621v2 Announce Type: replace-cross Abstract: Formal safety analysis determines whether a system admits a safe defense; adaptive evaluation characterizes the operating quality sustained under adversarial interaction. Both answers matter because systems with the same safety verdict can impose very different operational burdens. We introduce shielded analysis, a design-time framework that derives these answers from one encoded system while keeping the safety requirement and admissible threat model independently variable. It returns a defensibility certificate and a four-axis defensibility fingerprint spanning structural margin, shield latitude, and adaptive operating quality. Each axis is informative in its own right; their relationships show whether formal and operational assessments agree, diverge, or respond differently to system changes. We instantiate the framework for network defense on a reference segment and four controlled perturbations spanning topology, safety requirements, and adversary capabilities. Every configuration is certified defensible, yet two topology variants with nearly identical structural profiles sustain mean clean-host fractions of 22.7% and 80.7% under adaptive pressure. Shielded analysis turns a safety-game solution into a comparative instrument: it determines whether a defense exists, characterizes what that defense requires, and identifies which system changes strengthen it.
2026-09-16 04:00:00 · 招聘HR,榜单评测,论文
AI 资讯

Do LLMs Make Neural Distinguishers Wise?

arXiv cs.LGarXiv:2606.10692v2 Announce Type: replace-cross Abstract: Neural distinguishers are a cryptanalysis method for symmetric-key cryptography that trains machine learning models on pairs of plaintexts and ciphertexts with specific differences in order to recover a secret key. To the best of our knowledge, no existing work has explored the use of large language models (LLMs) for neural distinguishers. In this paper, we propose LLM-based neural distinguishers through a prompt design and conduct extensive experiments with them on SPECK-32/64 to investigate whether LLMs can strengthen neural distinguishers. We then found three key insights. First, by comparing the results of LLM-based neural distinguishers with ResNet in the existing work, we demonstrate that LLMs provide no observable improvement in the performance of neural distinguishers. Second, we confirm that, at high rounds, the choice of differences is no longer effective for LLM-based neural distinguishers as well as ResNet. Third, we show that the performance of LLM-based neural distinguishers can be significantly improved by incorporating only the XOR operation results as a prompt design.
2026-09-16 04:00:00 · 大模型,提示工程,招聘HR,论文
AI 资讯

Deep-learning-based low-energy trigger algorithms for the Hyper-Kamiokande experiment

arXiv cs.LGarXiv:2605.31391v2 Announce Type: replace-cross Abstract: Modern machine learning techniques have become increasingly important in particle physics because of their powerful pattern-recognition capabilities, including in real-time data acquisition where stringent runtime constraints apply. This paper details the performance of deep-learning-based trigger algorithms for a large water Cherenkov detector such as Hyper-Kamiokande, aimed at low-energy neutrino events (below 7 MeV). The performance of custom neural-network supervised classifiers is shown alongside two anomaly-detection approaches trained solely on detector noise: a pure autoencoder and a model based on Manifold Projection-Diffusion Recovery. The supervised model shows signal identification efficiencies of 76.7% for single electrons of 3 MeV kinetic energy, significantly exceeding signal efficiencies obtained from a traditional hit-count-based trigger of 26.4%, while the Manifold Projection-Diffusion Recovery approach reaches 35.4% at the same operating point. Runtime evaluations on GPU yield per-window inference latencies well below the millisecond scale.
2026-09-16 04:00:00 · 算力芯片,扩散模型,收购并购,论文
AI 资讯

Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks

arXiv cs.LGarXiv:2605.24696v3 Announce Type: replace-cross Abstract: Streaming intrusion-detection studies assemble evaluation streams from network captures by interleaving capture days, pooling captures, or replaying records round robin. We show on two benchmarks that this assembly is an uncontrolled experimental treatment changing what the evaluation measures. On CICIDS2017, reordering an identical record multiset under a fixed positional 70/15/15 split yields held-out samples sharing only 32.5% of their records, at prevalences of 68.235% and 25.2396% (42.9954 points apart), and reverses the measured ordering of the two deterministic scorers. Restricting both arms to the 78000 records both held out removes the reversal, so it is attributable to which records the assembly hands to the test set, not to the order in which the detector saw its history. That attribution assumes that history contributes no more on the records the arms do not share than on those they do. On LITNET-2020, pooling three temporally disjoint captures reports one 6.4982% operating point, the equal-weight mean of per-capture held-out prevalences from 0.176% to 15.7747%, an identity presented as an audit check. The evaluated detector's reset posterior P(r_t=0) equals the hazard rate exactly below the run-length cap, though evaluations spend nearly all their length at or beyond it, and its evaluated score is a function of P(r<=5), not of P(r=0). Its deployed max composition ranks worse than its tail term alone (0.103477 AP, 0.302658 AUC-ROC) because the auxiliary branch is inverted (AUC-ROC 0.281890) and the maximum lets it set the score wherever the tail is small. With evaluated records and fitted model fixed, changing only the accompanying batch moves the ECOD reference implementation's AUC-PR by 0.003063, so published ECOD numbers are not comparable across studies scoring different batches. Every measured value traces to an archived, hash-verified run manifest.
2026-09-16 04:00:00 · 扩散模型,模型评测,招聘HR,论文
AI 资讯

Subject-Specific Analysis of Self-Initiated Attention Shifts from EEG with Controlled Internal and External Attention Conditions

arXiv cs.LGarXiv:2605.18251v2 Announce Type: replace-cross Abstract: Self-initiated attention shifts play a critical role in voluntary behavior but are difficult to study due to the absence of explicit temporal markers. While previous studies have examined their neural correlates, it remains unclear how multi-dimensional electroencephalography (EEG) features contribute to their characterization within an interpretable computational framework. In this study, we build on an experimental paradigm developed in our previous work, which enables controlled comparison between task-constrained self-initiated shifts and externally instructed shifts under identical visual stimulation. Within this setting, we investigate whether preparatory EEG activity can distinguish these two types of attention shifts. We adopt a machine learning-based approach and conduct two complementary analyses: (1) a performance-oriented assessment of frequency-specific topographic patterns, and (2) a model-based feature attribution analysis using SHapley Additive exPlanations (SHAP). These analyses provide a structured view of how spectral features across regions of interest contribute to model behavior. Our results demonstrate reliable within-subject classification performance, indicating that preparatory EEG activity contains subject-specific discriminative information within this paradigm. The analysis shows that higher-frequency bands and frontal regions contribute strongly to model decisions, although such contributions should be interpreted cautiously due to the potential influence of non-neural artifacts in high-frequency EEG signals. Overall, this work highlights the value of interpretable machine learning for analyzing subject-specific EEG signal patterns in a controlled experimental setting, with potential applications in personalized and asynchronous brain-machine interface systems.
2026-09-16 04:00:00 · Transformer,扩散模型,招聘HR,榜单评测,论文
继续滚动加载更多…