Eventos
Histórias consolidadas
Uma história pode reunir várias matérias, fontes e sinais relacionados.
60 eventos exibidosLimpar filtros
Pesquisatrend 34
A promessa e o perigo de usar IA visual para estudar cidades
Em seu novo livro, “How AI Sees the City,” os líderes do Senseable City Lab do MIT examinam as implicações da tecnologia para a pesquisa da vida urbana.
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or earl…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Gap-free Differentially Private PCA for Gaussian Data
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
First-Order Stationarity of Reverse Diffusions
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consi…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Statistical attribute alignment for black-box generative AI via output post-processing
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we fu…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
User Model Extraction via Belief Self-Distillation
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
New LoRA Skills Should Read but Never Write
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes bot…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learnin…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the iss…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Trust Guided Decision Transformer
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a con…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Strategically Diverse Sampling for Self-Training
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
OC-GS: Gaussian Splatting for Irregular Turntable Capture
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores of 21.26, 19.36, and 15.83dB, respectively, exceeding all four evaluated pose-free Gaussian splatting baselines in each condition. Under a shared trainer, refining image-estimated angles improves mean foreground PSNR by 7.88dB over keeping those estimates fixed. An ablation study shows that both image-derived angle initialization and the shared motion model contribute to the improvement. On real…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attribut…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum
Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate for their students and classroom context. There is little known about how elementary teachers work as they implement conversational AI tools in real classrooms. In this study, we examine three teachers' experiences implementing an AI literacy and English Language Arts (ELA) curriculum built around ToyTalk, a conversational AI toy development platform, over 13 instructional days, a three-week summer camp. Drawing on daily individual reflections, group reflections, and post-camp interviews, we find that teachers' adaptive practices of repair, differentiation, translation, and balancing sit at the intersection of three tensions (techno…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese educ…
1 matériasRegulaçãoBR: Indireto
Pesquisatrend 34
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over whic…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this p…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three ha…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Two Conformal Constructions for Adaptive Within-Document AI-Text Screening
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal contro…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
A Flow Matching Framework for Neural Representational Dissimilarity
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different velocity constraints. We find that flow matching has advantages for estimating distances involving complicated distributions and continuous variables. Furthermore, this framework enables the design of new distance metrics in a principled way. Together, flow matching provides a unified approach for understanding, estimating, and designing neural representational dissimilarity metrics.
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Can You Check That? The Checkability Boundary for Local LLM Network Automation
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while esc…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-con…
1 matériasPesquisaBR: Referência estratégica
Pesquisatrend 34
Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.94B parameters) and a Modern Standard Arabic TTS model (Fasih-TTS-V1) that ranks 5th of 17 overall and 2nd of 11 open-weight systems on the community-voted Arabic TTS Arena for MSA. Second, an account and metering layer - a free per-account turn allowance, capacity-aware refusal, and email verification deferred to the point it actually matters - that turns an open demo into an operable, abuse-resi…
1 matériasLançamentoBR: Referência estratégica
Pesquisatrend 34
ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ClearGS further introduces Render-Guided In-Video Restoration (RIVR). The current 3DGS render provides a pose-aligned structural candidate, a frozen no-reference restoration expert restores the corresponding raw video observation without any clean reference image, and no-reference perceptual scores select among the render, restored observation, and high-frequency fused candidate. ClearGS then app…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Evaluating Cultural Awareness of LLMs for Haitian Creole
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What thos…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performance differences across adaptation strategies, while model scaling yields varying detection gains across datasets. We further observe that models with comparable detection accuracy can exhibit markedly different computational costs, and that low-bit quantization largely preserves detection performance in the evaluated configurations. Finally, we examine detection robustness under structural, semanti…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Equation discovery with Bayesian tree-adjoining grammars
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, and a Reversible-Jump MCMC sampler with structure-preserving tree moves is used to infer the joint posterior over model structure, parameters and predictions. Two training objectives are considered; that is, a one-step-ahead objective with conjugate parameter proposals, and a simulation-based objective handled by likelihood-free inference. The approach is validated on a simulated polynomial NARX …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel g…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Open Vocabulary Domain Unlearning
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Dom…
1 matériasRegulaçãoBR: Não identificado
Pesquisatrend 34
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that expo…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent blueprint that bridges an agent's body and brain. Through AdaConcat, Morphogene jointly conditions morphology and control generation at the limb level, allowing its variations to induce coordinated changes in both components. Building on this representation, we propose \textbf{GeCode}, which formulates co-design as exploration in the compact Morphogene space. Each Morphogene anchors a local design region in which nearby body--brain designs are explored, wh…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process without necessarily changing the dynamics to be learned. We propose STFO (Spatio-Temporal Field Operator), which parameterizes forecasting knowledge as a shared field-evolution operator and handles changing sensor layouts through observation and query interfaces. Normalized coordinate-based aggregation lifts irregular sensor histories onto a fixed latent grid, enabling reuse of learned spatial maps across observation sets without sensor-specific param…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Marčenko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID attenuates factor-dominated variation and recovers contemporaneous (lag-$0$) structure from the resulting innovations, with edge selection calibrated against a data-driven edge-free null. Rather than being tied to a particular discovery algorithm, it can wrap existing discovery engines; we demonstrate consistent improvements across three such methods. On a diverse synthetic out-of-distribution benchmark …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Benchmarking Attention for Tabular Foundation Models
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Geometric Moment Contraction for Stochastic Nesterov Acceleration
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=Θ_k+β(Θ_k-Θ_{k-1}),\qquad Θ_{k+1}=Y_k-γG(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $βγL_p
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Agentes LLM podem facilmente adulterar seus próprios rastros
Monitoramento assíncrono, investigações de incidentes e auditorias de conformidade dependem principalmente dos rastros de agentes para reconstruir o que ocorreu. Essas análises presumem que agentes LLM não podem adulterar seus próprios rastros de execução. Demonstramos que agentes LLM locais como Claude Code, Codex, Antigravity, Open Code e Grok Build não conseguem impor essa barreira. Todos os harnesses testados, exceto o Muse Code, permitiram que os agentes excluíssem seus rastros quando solicitados, sem acionar as salvaguardas do monitor. Também validamos que invasores externos podem explorar essa lacuna para induzir a exclusão de rastros. Por fim, mostramos que o comportamento de adulteração de rastros surge naturalmente em modelos de ponta, quando os agentes buscam melhorar suas recompensas. Recomendamos que os profissionais garantam que o registro de rastros ocorra por meio de um mecanismo de interceptação independente, fora do controle do agente, preservando a integridade dos rastros mesmo em casos de comprometimento total do host. No geral, o…
1 matériasIncidenteBR: Não identificado
Pesquisatrend 34
AD-WM: Modelos de Mundo Discriminativos por Ação para Controle Preditivo de Modelo Contrafactual
Modelos de mundo latentes são tipicamente treinados para prever transições factuais, enquanto o controle preditivo de modelo (MPC) deve comparar ações alternativas a partir do mesmo estado. Assim, um modelo pode alcançar baixo erro de previsão factual, mas distinguir mal as ações candidatas. Introduzimos o AD-WM, um modelo de mundo de incorporação conjunta discriminativo por ação para MPC contrafactual. O AD-WM combina dinâmicas latentes residuais com regularização de recuperação de ação em nível de preditor, usando dinâmica inversa e um objetivo de recuperação normalizado motivado por informação mútua condicional. Ambos os objetivos incentivam que as transições de planejamento preservem a informação da ação; suas cabeças auxiliares são descartadas no tempo de teste, mantendo o MPC inalterado. No OGBench-Cube, o AD-WM eleva o sucesso de partida difícil de 3,7% para 52,0% em relação a uma linha de base LeWM correspondente e melhora o sucesso médio sobre a linha de base reproduzida em quatro de cinco ambientes de simulação…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Inversão de Gradiente Temporal para Reconstrução Privada de Trajetórias em Aprendizado por Reforço Incorporado
Aprendizado distribuído em agentes de aprendizado por reforço incorporado oferece um grau de privacidade ao manter os dados brutos dos sensores no dispositivo e transmitir apenas gradientes de política ao servidor. Contudo, a estrutura temporal pode amplificar esse vazamento além de ataques de quadro único. Introduzimos o Temporal Reconstruction Attack on Consecutive Encodings (TRACE), um ataque amortizado de inversão de gradiente temporal que reconstrói autoregressivamente a sequência de trajetórias privadas de observação‑ação a partir dos gradientes de aprendizado de política por passo. O ataque explora dois sinais estruturais ignorados por métodos de quadro único anteriores: (i) correlação intertemporal entre gradientes incorporados sucessivos, que formalizamos por meio de um limite de informação mútua condicional, e (ii) recuperação em forma fechada de ações a partir da estrutura de gradiente da cabeça de política, que provamos ser exata quando a regularização de entropia padrão é suficientemente pequena. Em dados reservados …
1 matériasIncidenteBR: Não identificado
Pesquisatrend 34
Detecção Agente de Conspirações Online
Discurso conspiratório nas redes sociais nem sempre é expresso por meio de afirmações explícitas ou marcadores lexicais estáveis. O mesmo conteúdo superficial pode expressar apoio, preocupações legítimas, crítica, sátira ou zombaria. O principal desafio, portanto, não é apenas reconhecer afirmações relacionadas a conspirações, mas inferir a intenção do falante — a força ilocutória da enunciação. Argumentamos que isso pode ser alcançado mediante o uso de contextos sociais relevantes e propomos uma estrutura agente, equipada com um conjunto de ferramentas que suportam consultas sociais. Demonstramos os benefícios de nossa abordagem em um conjunto de dados único de tweets em hebraico, cobrindo 80 %–90 % dos tweets públicos em hebraico publicados ao longo de quatro anos (final de 2018 – início de 2023), abrangendo vários ciclos eleitorais, bem como os anos da pandemia de COVID e campanhas de vacinação relacionadas. Essa cobertura extensiva pode ser usada na recuperação de diferentes …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
RAPID: Programação Agente de Robô a partir de Demonstrações
Agentes de codificação demonstraram enorme sucesso na resolução de problemas de programação complexos. Para aproveitar seu potencial em sistemas robóticos, este trabalho apresenta Robot Agentic Programming from Demonstrations (RAPID), que gera, verifica e refina automaticamente programas de robô a partir de uma única demonstração visual humana. O loop agente iterativo de refinamento de código requer vários componentes essenciais: (i) uma especificação de tarefa testável, (ii) primitivas de ação para execução do robô e (iii) um ambiente interativo para execução e verificação do programa. O RAPID infere automaticamente os três a partir da demonstração. Para tornar o programa resultante reutilizável além do cenário de demonstração, o RAPID utiliza uma representação de programa relacional centrada em objetos que foca na estrutura subjacente da estratégia demonstrada em vez do movimento específico em si: expressa as primitivas de ação…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Rolling-WAM: Modelos de Ação Mundial com Imaginação Contínua
Modelos de Ação Mundial (WAMs) combinam geração de ação com previsão visual futura para manipulação robótica. Contudo, concluir o processo conjunto de desnoising de vídeo e ação a cada ciclo de replanejamento gera latência substancial, atrasando atualizações de ação e limitando a responsividade em loop fechado. Apresentamos o Rolling-WAM, uma formulação que distribui o desnoising conjunto ao longo de ciclos de replanejamento sucessivos. Nosso método mantém uma janela deslizante de blocos de vídeo-ação em níveis de ruído escalonados. Em cada etapa, uma agenda de ruído contínuo desnoisa totalmente o bloco de ação iminente para execução, enquanto refina parcialmente blocos de futuro mais distante. À medida que a janela avança com novas observações de câmera, os blocos futuros retidos continuam seu processo de desnoising. Isso distribui o custo computacional ao longo do tempo, mantendo um contexto visual-ação evolutivo entre os limites dos blocos. Avaliações em LIBERO, RoboTwin, um…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
JevOut: Contexto Natural Pode Inverter Modelos de Decisão
Modelos de decisão dedicados, como o Jev, mapeiam linguagem não estruturada para distribuições de probabilidade sobre escolhas finitas, permitindo que suas saídas roteiem solicitações, selecionem ferramentas e acionem ações diretamente. Contudo, entradas do mundo real raramente chegam isoladamente: vêm acompanhadas de detalhes de fundo e contexto circundante. Descobrimos que pequenas adições que se encaixam naturalmente nesse contexto podem, no entanto, redirecionar uma decisão que seria correta, mesmo quando a resposta correta permanece inalterada. Para estudar esse comportamento, fixamos uma opção alvo incorreta para cada item inicialmente correto e usamos as probabilidades das opções do modelo para refinar adições contextuais fluentes, preservando a fonte, a pergunta, as escolhas e a resposta de referência. Em 64 avaliações de alvo aceitas, o otimizador identifica contextos que redirecionam o Jev em 312 de 508 decisões inicialmente corretas (61,4 %); em 229 casos, o Jev atribui probabilidade de pelo menos 0,7 …
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
SemMSA: Análise de Sentimento Multimodal Robusta Auxiliada por Semântica Latente com Dados Incompletos
Pesquisas recentes em Análise de Sentimento Multimodal (MSA) têm se concentrado em aprender a partir das modalidades de linguagem, visual e acústica com dados incompletos para inferir o sentimento humano. A maioria dos estudos costuma compensar a informação ausente reconstruindo características das modalidades ou projetando mecanismos de fusão complexos. Contudo, esses métodos ainda sofrem com geração espúria e orientação ruidosa devido à falta de fundamentação semântica de alto nível em evidências multimodais parcialmente observadas. Para enfrentar esses problemas, propomos o SemMSA, uma estrutura auxiliada por semântica latente que constrói semânticas ricas relevantes ao sentimento usando LLMs, integrando totalmente todas as modalidades via alinhamento espectral sem âncora. Ela consiste principalmente em Refinamento Semântico Cross-modal (CSR) e Alinhamento Espectral Cross-modal (CSA). Especificamente, o CSR primeiro extrai adaptativamente representações visuais e acústicas por meio de …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Agentes de Codificação para Problemas Generalizados de Planejamento de Tarefas e Movimento
Problemas de planejamento de tarefa e movimento (TAMP) permanecem difíceis mesmo com observabilidade total e estados centrados em objetos, pois decisões discretas estão intimamente acopladas a restrições geométricas, cinemáticas e dinâmicas. O TAMP generalizado aborda essa dificuldade explorando regularidades entre instâncias de problema para reduzir o esforço de planejamento em novas instâncias. Contudo, os métodos existentes exigem engenharia substancial específica de TAMP. Investigamos se agentes de codificação podem automatizar esse processo sintetizando programas que se generalizem entre instâncias. Dada uma descrição de tarefa e acesso a um simulador, cada agente escolhe como interagir com o ambiente enquanto desenvolve um programa dentro de um orçamento fixo de síntese. O programa é então congelado e avaliado em instâncias não vistas. Avaliamos Claude Code (Opus 5) e Codex (GPT-5.6 Sol e GPT-6 Astra) em 28 ambientes simulados de KinDER e PDDLStream, …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Confiar ou Não Confiar: Verificação de Fatos Aumentada por Recuperação em Discurso
A desinformação online aparece cada vez mais em formatos falados, como clipes de notícias, podcasts, entrevistas, discursos políticos e vídeos em redes sociais, gerando a necessidade de sistemas de verificação de fatos que possam validar afirmações diretamente a partir da fala. Introduzimos o VeriSpeak, um benchmark de sondagem para estudar a verificação de fatos baseada em fala em Large Audio Language Models (LALMs). O VeriSpeak contém 3.879 afirmações faladas abrangendo fatos temporais, geográficos e relacionais, com rótulos equilibrados de verdadeiro e falso. O benchmark foi projetado para examinar se a capacidade de verificação factual transfere do texto para a fala, e se LALMs aumentados por recuperação podem usar evidências textuais para apoiar ou refutar corretamente afirmações faladas. Nossos experimentos revelam uma lacuna consistente entre as modalidades texto‑fala: LALMs que verificam afirmações escritas de forma confiável frequentemente falham nas mesmas afirmações quando faladas. Além disso, a recuperação isolada fornece ganho limitado …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
PoEM: Predizendo Resultados de RL a partir de Políticas Existentes
Modelos fundacionais são pós-treinados com reinforcement learning (RL) para maximizar recompensas específicas, como alinhamento humano, correção ou seguimento de instruções. Esse processo de pós‑treinamento é computacionalmente intensivo, às vezes instável, e precisa ser reiniciado do zero sempre que o modelo de recompensa muda ou quando queremos combinar múltiplas recompensas. Assim, perguntamos: dado uma nova função de recompensa, é possível prever os resultados de RL sem realmente executar RL sobre ela? Respondemos afirmativamente ao introduzir o PoEM, uma estrutura para prever as saídas de RL em uma nova função de recompensa usando um conjunto de modelos já pós‑treinados em outras recompensas. Primeiro, demonstramos que se a nova função de recompensa pode ser escrita como uma combinação linear das existentes, então a nova política em log‑space pode ser escrita como uma combinação linear das políticas log‑existentes. Surpreendentemente, mesmo em casos onde a re…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-vers…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were …
1 matériasRegulaçãoBR: Não identificado
Pesquisatrend 34
Direcionamento Minimante Invasivo de Modelos de Linguagem
Direcionamento pré‑logit adapta um modelo de linguagem congelado a uma recompensa no momento do teste ao adicionar vetores aos seus estados ocultos finais. A otimização de recompensa não regularizada pode alterar substancialmente a distribuição de saída e degradar a qualidade da geração. Propomos o Minimally Invasive Steering Vector Optimization (MISVO), que penaliza intervenções usando a geometria KL local da distribuição de tokens induzida. O quadrático de Fisher resultante mede a sensibilidade distributiva e admite um gradiente analítico computado por meio de produtos matriz‑vetor com a cabeça do modelo de linguagem congelado. Derivamos uma decomposição exata do gradiente KL ao nível de sequência em um termo Fisher analítico e um termo de função‑pontuação de sufixo. Para um horizonte de geração fixo, mostramos que o termo de sufixo é de segunda ordem na magnitude do direcionamento e que três substitutos de Fisher concordam com o gradiente KL completo na primeira ordem. O MISVO usa o …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Um Limite Inferior Quase Quadrático para Otimização Linear sobre Corpos Convexos no Modelo de Oráculo de Pertencimento
Demonstramos limites inferiores quase quadráticos para algoritmos aleatórios de otimização linear e amostragem uniforme sobre corpos convexos no modelo de oráculo de pertencimento. Para otimização linear, isso corresponde ao limite superior quase quadrático conhecido, até um fator poli‑logarítmico na dimensão. Para amostragem uniforme, isso melhora o limite inferior linear anterior. Nossa construção também implica o mesmo limite inferior para estimativa de volume.
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Métodos Extra-Proximais Ancorados: Métodos Ótimos de Ordem Superior para Problemas de Inclusão Monótona
Estudamos a complexidade determinística de oráculo para encontrar soluções aproximadas de problemas de inclusão monótona compostos, formados pela soma de um operador monótono suave de valor único e um operador monótono maximamente multivalorado, sob o critério de resíduo tangente. Introduzimos a estrutura Anchored Extra-Proximal (AEP), que combina um passo de extrapolação ancorado com uma atualização proximal ancorada imprecisa que satisfaz uma condição de erro relativo. A estrutura recupera o método Fast Extragradient composto no cenário de primeira ordem e gera extensões naturais de segunda e ordem superior ao substituir o operador na atualização implícita por sua aproximação de Taylor no ponto extrapolado. Para todo $p\geq 2$, assumindo que a derivada $(p-1)$‑ésima do operador de valor único é Lipschitz contínua, combinamos essa construção com uma busca de linha por bissecção para obter um método de ordem $p$…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
The Alignment Illusion in Multimodal Large Language Models
Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common o…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbo…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers
Latent neural surrogate solvers, or latent dynamics models, accelerate simulations of time-dependent physical systems by evolving a compressed latent space rather than resolving full-resolution fields directly. In principle this reduces computational cost and simplifies learning, but in practice errors often accumulate rapidly during long autoregressive rollouts, limiting predictive utility. We show that this instability does not stem from the latent representation itself, but arises when it is trained solely for reconstruction, producing representations poorly suited to long-horizon forecasting. We systematically evaluate training-level interventions that align latent representations with long-horizon rollout: Koopman operator learning and Hamming noise injection during autoencoder training to improve compression, together with noise injection and multi-step rollout fine-tuning to impr…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Intrinsic-Extrinsic Coupling in Learning Dynamics
A learner's current observations need not determine its response to further training. We formulate intrinsic-extrinsic coupling through the continuation-conditioned value of a constrained learning-state intervention, with observation-relative fibers describing present agreement. An executable finite-frame classifier-head write protects current logits while repairing specified historical margins under finite-precision acceptance checks. We distinguish local admissibility, continuation-conditioned intervention value, and complete-policy performance. A matched four-cell contrast identifies readout-specific non-additivity between the same intrinsic intervention and alternative external continuations. In a CLINC-derived class-incremental setting, replay changes the write's 32-update contribution from five correct predictions to zero. Nonzero interactions also occur under output distillation,…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints
U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct document-level Event Knowledge Graphs (EKGs) from CourtListener complaints. ARGUS extracts fact-bearing statements, builds chunk-level event graphs with participant, temporal, and causal structure, and merges them into document-level representations. We evaluate graph quality through human and multi-model assessment and test downstream utility on claim classification and legal QA. The graph-structured classifier outperforms raw and linearized baselines on the held-out set, and EKG-only retrieval improves document-scoped QA, while open-retrieval gains remain limited by low first-…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Do Audio Language Models Hear and Read Distinctive Features Alike?
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across …
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition
Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation o…
1 matériasPesquisaBR: Não identificado
Pesquisatrend 34
Does a model's stated reason for rejecting a candidate do any work?
Asked to choose between candidates and explain the choice, a language model often rejects a rival by naming a fact its profile lacks: no director, no date of death. That sentence is a claim about the text in front of the model, and it can be tested without any judge. We insert a real corpus sentence stating the named fact into the rival's profile and ask again under greedy decoding. Two controls separate content from placement: a length-matched irrelevant sentence at the same profile, and the same two sentences at a third option the model never mentioned. In the largest of three runs, six open models on 2WikiMultihopQA, supplying the named fact at the profile the model named moves its choice more than the irrelevant control does, odds ratio 3.57 [1.54, 8.26], Holm p=0.0210, and this survives dropping any single model. The contrast the design was built to detect, the same fact at the opt…
1 matériasPesquisaBR: Não identificado
Prefere começar pelo PDF de hoje? Baixe o panorama.