Eventos
Histórias consolidadas
Uma história pode reunir várias matérias, fontes e sinais relacionados.
60 eventos exibidosLimpar filtros
Chips e hardwaretrend 100
Um investimento de US$ 1.000 dividido entre Alphabet e Nvidia valerá isso até 2030 - fool.com
Um investimento de US$ 1.000 dividido entre Alphabet e Nvidia valerá isso até 2030 fool.com
3 matériasInvestimentoBR: Referência estratégica
Modelos e laboratoriostrend 94
Got $5,000?: 2 Magnificent AI Cloud Stocks to Buy Before the Anthropic IPO - fool.com
Got $5,000?: 2 Magnificent AI Cloud Stocks to Buy Before the Anthropic IPO fool.com
3 matériasInvestimentoBR: Indireto
Chips e hardwaretrend 94
Nvidia’s Stock Is Flashing a Warning Sign as Valuation Drops - Bloomberg.com
Nvidia’s Stock Is Flashing a Warning Sign as Valuation Drops Bloomberg.com
3 matériasInvestimentoBR: Referência estratégica
Startups e investimentostrend 90
Itaú anuncia investimento na Simile, startup de IA de US$ 2 bilhões que simula comportamento humano - timesbrasil.com.br
Itaú anuncia investimento na Simile, startup de IA de US$ 2 bilhões que simula comportamento humano timesbrasil.com.br
3 matériasInvestimentoBR: Direto
Chips e hardwaretrend 76
3 razões pelas quais a Nvidia se encaixa no estilo de investimento de Warren Buffett - Yahoo Finance
3 razões pelas quais a Nvidia se encaixa no estilo de investimento de Warren Buffett Yahoo Finance
2 matériasInvestimentoBR: Referência estratégica
Startups e investimentostrend 68
Rodadas da semana: Fever capta US$ 250 milhões, e Itaú aposta em startup de IA dos EUA - bloomberglinea.com.br
Rodadas da semana: Fever capta US$ 250 milhões, e Itaú aposta em startup de IA dos EUA bloomberglinea.com.br
2 matériasInvestimentoBR: Direto
Startups e investimentostrend 64
Uma análise levava duas semanas. Com agentes de IA, a VOX diz fazer em seis horas - businessmoment.com.br
Uma análise levava duas semanas. Com agentes de IA, a VOX diz fazer em seis horas businessmoment.com.br
2 matériasInvestimentoBR: Direto
Startups e investimentostrend 64
Ex-CFO do Nubank vira sócio da Kaszek, que capta novo fundo - neofeed.com.br
Ex-CFO do Nubank vira sócio da Kaszek, que capta novo fundo neofeed.com.br
2 matériasInvestimentoBR: Direto
Startups e investimentostrend 64
A nova fase da corrida da inteligência artificial - estadao.com.br
A nova fase da corrida da inteligência artificial estadao.com.br
2 matériasInvestimentoBR: Direto
Startups e investimentostrend 64
Itaú investe em Simile, startup de IA dos EUA - aciara.com.br
Itaú investe em Simile, startup de IA dos EUA aciara.com.br
2 matériasInvestimentoBR: Direto
Nuvem e infraestruturatrend 64
Mercado para empresas de data centers que buscam IPOs fica mais desafiador - CNN Brasil
Mercado para empresas de data centers que buscam IPOs fica mais desafiador CNN Brasil
2 matériasInvestimentoBR: Direto
Nuvem e infraestruturatrend 64
Sanção do Redata reduz custo de data centers e amplia espaço para investimentos no Brasil - Times Brasil | CNBC
Sanção do Redata reduz custo de data centers e amplia espaço para investimentos no Brasil Times Brasil | CNBC
2 matériasInvestimentoBR: Direto
Modelos e laboratoriostrend 59
Fonte: provedor de inferência Modal Labs se aproximando de rodada de US$ 750 mi com avaliação de US$ 15,75 bi
A nova captação deve mais que triplicar a avaliação da startup de infraestrutura de IA em relação a apenas quatro meses atrás.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 56
Mavi aposta no boom da IA criando demanda por um novo tipo de contador
A empresa de recrutamento contábil Mavi sai do modo stealth com US$ 4 milhões em financiamento.
1 matériasInvestimentoBR: Não identificado
Brasil e estrategiatrend 56
Brasil | Cade arquiva investigação sobre investimentos da Amazon na Anthropic - dplnews.com
Brasil | Cade arquiva investigação sobre investimentos da Amazon na Anthropic dplnews.com
9 matériasInvestimentoBR: Direto
Startups e investimentostrend 55
Esta investidora precoce da Groq espera que metade de suas apostas falhe
Em entrevista ao Crunchbase News, Sandhya Venkatachalam, fundadora e sócia‑gerente da Axiom Partners, discute por que ela vai além de perfis de fundadores familiares, o que torna uma empresa de IA durável e como um investimento precoce na Groq moldou sua abordagem.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 50
Mercado tem US$ 287 milhões em movimento com fintechs, IA e entretenimento na semana - Startupi
Mercado tem US$ 287 milhões em movimento com fintechs, IA e entretenimento na semana Startupi
1 matériasInvestimentoBR: Direto
Startups e investimentostrend 41
Startups: VOX lança novo fundo e avança na tese de AI-first dentro de casa - infomoney.com.br
Startups: VOX lança novo fundo e avança na tese de AI-first dentro de casa infomoney.com.br
3 matériasInvestimentoBR: Direto
Startups e investimentostrend 39
Criei um avatar digital interativo de mim mesmo — e você pode conversar com ele
Depois de obter um avatar interativo e treiná‑lo para discutir fraudes em venture, tenho sentimentos ambíguos sobre criar clones de IA de nós mesmos.
1 matériasInvestimentoBR: Não identificado
Chips e hardwaretrend 39
Nvidia negocia próximo ao seu recorde de 52 semanas com a avaliação mais barata em uma década. A história indica que isso pode valer US$ 1.000 investidos até 2030. - Yahoo Finance
Nvidia negocia próximo ao seu recorde de 52 semanas com a avaliação mais barata em uma década. A história indica que isso pode valer US$ 1.000 investidos até 2030. Yahoo Finance
1 matériasInvestimentoBR: Referência estratégica
Startups e investimentostrend 34
Abvcap Experience 2026 debate IA, expansão global, search funds e rotas de saída no capital privado - Pequenas Empresas & Grandes Negócios
Abvcap Experience 2026 debate IA, expansão global, search funds e rotas de saída no capital privado Pequenas Empresas & Grandes Negócios
1 matériasInvestimentoBR: Direto
Startups e investimentostrend 34
Modelos chineses de código aberto podem ajudar Brasil a acelerar na corrida global da IA, dizem investidores - Época Negócios
Modelos chineses de código aberto podem ajudar Brasil a acelerar na corrida global da IA, dizem investidores Época Negócios
1 matériasInvestimentoBR: Direto
Startups e investimentostrend 34
Itaú investe em IA da Simile para simular cenários e decisões - IT Forum
Itaú investe em IA da Simile para simular cenários e decisões IT Forum
1 matériasInvestimentoBR: Direto
Startups e investimentostrend 34
Firecrawl capta US$ 75 milhões para organizar conhecimento para agentes de IA - startups.com.br
Firecrawl capta US$ 75 milhões para organizar conhecimento para agentes de IA startups.com.br
1 matériasInvestimentoBR: Direto
Modelos e laboratoriostrend 34
Antes da IPO nos EUA, a neocloud britânica de IA Nscale garante US$ 3,36 bilhão em financiamento conversível.
O financiamento, que vem da Third Point, Nvidia e outros, vai impulsionar a construção massiva de data centers de IA da empresa.
1 matériasInvestimentoBR: Indireto
Modelos e laboratoriostrend 34
Fundadores da Anthropic buscam controle de voto antes da IPO
Anthropic está pedindo aos seus acionistas que aprovem uma estrutura que concederia aos seus sete cofundadores um total combinado de 50,1% dos votos na maioria dos assuntos corporativos.
1 matériasInvestimentoBR: Referência estratégica
Modelos e laboratoriostrend 34
CEO da ElevenLabs sobre margens, cronograma de IPO e informar aos clientes que estão conversando com um bot
ElevenLabs fornece a voz de IA no outro lado de muitas chamadas de atendimento ao cliente, e seu CEO me disse esta semana que as empresas provavelmente deveriam lhe contar isso — pelo menos até que obter uma máquina seja o que todos esperam, de qualquer forma.
1 matériasInvestimentoBR: Não identificado
Modelos e laboratoriostrend 34
Ando quer desafiar o Slack com um aplicativo de mensagens para equipes que permite que humanos e agentes trabalhem juntos
Ando levantou US$ 20 mi em financiamento pré-seed e seed de investidores como Accel, Index Ventures e Emergence.
1 matériasInvestimentoBR: Não identificado
Modelos e laboratoriostrend 34
Snorkel AI triplica avaliação para $3.5B à medida que a demanda por dados de treinamento de IA dispara
A startup de sete anos levantou $350 million Series E para impulsionar sua abordagem de data-as-a-service.
1 matériasInvestimentoBR: Não identificado
Modelos e laboratoriostrend 34
O IPO da Nscale testará novamente o apetite de Wall Street por apostas concentradas em IA.
O desenvolvedor britânico de data centers de IA depende das gigantes de tecnologia Microsoft e Anthropic pela maior parte de sua receita.
1 matériasInvestimentoBR: Indireto
Startups e investimentostrend 34
Lightspeed mira US$ 250 milhões para novo fundo na Índia, focado em IA em estágio inicial
A empresa do Vale do Silício está alinhando seu ciclo de captação de recursos na Índia com seus fundos globais pela primeira vez, ao mudar para um período de investimento mais curto.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
Shield AI, Waabi e General Motors sobre construir IA quando o fracasso não é uma opção no TechCrunch Disrupt 2026
Líderes da Waabi, Shield AI e General Motors se juntam ao Palco Real World AI no TechCrunch Disrupt 2026 para falar sobre construção de IA. Economize até US$ 200 até 25 de setembro, 23h59 PT. Adquira um segundo passe com 50 % de desconto.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
Ema arrecada US$ 77 milhões enquanto a IA começa a impactar softwares e serviços empresariais
Ema já arrecadou US$ 140 milhões até o momento e possui mais de 50 clientes corporativos, incluindo Google e Microsoft.
1 matériasInvestimentoBR: Indireto
Produtos e aplicacoestrend 34
Andreessen Horowitz está lançando uma ‘academia’ sem dever de casa e com parcerias com Palantir, Google e Meta
A firma de capital de risco Andreessen Horowitz (a16z) está criando uma "academia" posicionada como um canal para jovens construírem ou ingressarem em uma startup do Vale do Silício. A "Horowitz Andreessen Academy" será lançada com 10 parceiros, incluindo Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit e Stripe, além de US$ 42 milhões em financiamento liderado pela a16z. […]
1 matériasInvestimentoBR: Indireto
Startups e investimentostrend 34
Cinco sessões de segurança em IA que todo fundador deve incluir na agenda do TechCrunch Disrupt 2026
No TechCrunch Disrupt 2026, cinco sessões nos palcos AI Stage e Real World AI Stage abordam a segurança em IA, com líderes da Anthropic, NVIDIA, AWS, Waabi e outros. Inscreva‑se agora e economize até $200 antes de 25 de set.
1 matériasInvestimentoBR: Indireto
Startups e investimentostrend 34
Morphotonics arrecada €40 mi para expandir sua tecnologia de exibição para data centers
Empresa deeptech Morphotonics arrecada €40 mi de investidores, incluindo 3M Ventures, Innovation Industries, BOM e Invest‑NL.
1 matériasInvestimentoBR: Referência estratégica
Modelos e laboratoriostrend 34
Como o acordo de IA da Anthropic na Akamai Technologies (AKAM) mudou sua história de investimento - simplywall.st
Como o acordo de IA da Anthropic na Akamai Technologies (AKAM) mudou sua história de investimento simplywall.st
1 matériasInvestimentoBR: Referência estratégica
Modelos e laboratoriostrend 34
Anthropic lança modelo de IA mais barato antes do IPO - Financial Times
Anthropic lança modelo de IA mais barato antes do IPO Financial Times
1 matériasInvestimentoBR: Referência estratégica
Startups e investimentostrend 34
Os 10 maiores rodadas de financiamento da semana: Cibersegurança, IA e Saúde lideram
Esta semana trouxe uma abundante oferta de grandes rodadas de financiamento de startups, liderada por dois financiamentos de $400 million para unicórnios de cibersegurança e grandes financiamentos para startups em setores em alta, incluindo IA fundamental, descoberta de medicamentos, neurotecnologia e até rainmaking.
1 matériasInvestimentoBR: Referência estratégica
Startups e investimentostrend 34
Demissões no setor de tecnologia superam 2025 enquanto grandes empresas mudam gastos para IA
De janeiro a agosto, as demissões no setor de tecnologia dos EUA chegaram a pelo menos 94.046, um aumento de 16,8% em relação a 80.486 no mesmo período de 2025. Curiosamente, e como era de se esperar, muitas das demissões ocorreram enquanto as empresas de tecnologia redirecionavam gastos para IA e reestruturavam operações para reduzir custos.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
Exclusivo: de ligações de reserva a check-ins tardios, Dextr AI arrecada US$ 6,7 mi para agentes de IA em hotéis
Dextr AI está saindo do modo stealth com US$ 6,7 milhões em financiamento semente para criar agentes que gerenciam reservas, solicitações de hóspedes, coordenação de equipe e outras tarefas hoteleiras.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
O mapa emergente de fusões e aquisições para segurança de agentes de IA
À medida que agentes de IA obtêm acesso a dados, sistemas e ferramentas empresariais, eles surgem como uma nova classe de identidade ativa que requer permissões especializadas, monitoramento e governança, escreve o autor convidado Itay Sagie. Ele acredita que o mercado, e suas oportunidades de fusões e aquisições, provavelmente se formarão em torno de pontos de controle específicos, tornando o posicionamento preciso crucial para startups.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
Exclusivo: Você pode confiar naquele agente de IA? Baselayer levanta US$ 35 milhões para ajudar empresas a decidir
Baselayer, uma startup impulsionada por IA que ajuda instituições financeiras a verificar empresas e avaliar risco de fraude, levantou US$ 35 milhões em Série A liderada pela M13 para expandir sua tecnologia de identidade para agentes de IA.
1 matériasInvestimentoBR: Não identificado
Startups e investimentostrend 34
À medida que VCs de software perseguem ex‑funcionários da SpaceX, um veterano de tecnologia de defesa alerta sobre ‘turistas e FOMO’
Em entrevista ao Crunchbase News, Van Espahbodi, sócio geral da Generational Partners, discute como a IA está mudando a economia de hardware, por que investidores de software estão correndo para a tecnologia industrial e o que ele acredita que muitos deles não compreendem sobre o setor.
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the iss…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three ha…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Evaluating Cultural Awareness of LLMs for Haitian Creole
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Rolling-WAM: Modelos de Ação Mundial com Imaginação Contínua
Modelos de Ação Mundial (WAMs) combinam geração de ação com previsão visual futura para manipulação robótica. Contudo, concluir o processo conjunto de desnoising de vídeo e ação a cada ciclo de replanejamento gera latência substancial, atrasando atualizações de ação e limitando a responsividade em loop fechado. Apresentamos o Rolling-WAM, uma formulação que distribui o desnoising conjunto ao longo de ciclos de replanejamento sucessivos. Nosso método mantém uma janela deslizante de blocos de vídeo-ação em níveis de ruído escalonados. Em cada etapa, uma agenda de ruído contínuo desnoisa totalmente o bloco de ação iminente para execução, enquanto refina parcialmente blocos de futuro mais distante. À medida que a janela avança com novas observações de câmera, os blocos futuros retidos continuam seu processo de desnoising. Isso distribui o custo computacional ao longo do tempo, mantendo um contexto visual-ação evolutivo entre os limites dos blocos. Avaliações em LIBERO, RoboTwin, um…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
JevOut: Contexto Natural Pode Inverter Modelos de Decisão
Modelos de decisão dedicados, como o Jev, mapeiam linguagem não estruturada para distribuições de probabilidade sobre escolhas finitas, permitindo que suas saídas roteiem solicitações, selecionem ferramentas e acionem ações diretamente. Contudo, entradas do mundo real raramente chegam isoladamente: vêm acompanhadas de detalhes de fundo e contexto circundante. Descobrimos que pequenas adições que se encaixam naturalmente nesse contexto podem, no entanto, redirecionar uma decisão que seria correta, mesmo quando a resposta correta permanece inalterada. Para estudar esse comportamento, fixamos uma opção alvo incorreta para cada item inicialmente correto e usamos as probabilidades das opções do modelo para refinar adições contextuais fluentes, preservando a fonte, a pergunta, as escolhas e a resposta de referência. Em 64 avaliações de alvo aceitas, o otimizador identifica contextos que redirecionam o Jev em 312 de 508 decisões inicialmente corretas (61,4 %); em 229 casos, o Jev atribui probabilidade de pelo menos 0,7 …
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search
Search quality evaluation provides essential supervision and diagnostic signals for the development and iteration of industrial search systems. Although large language models (LLMs) offer a scalable alternative to manual assessment, reliable automatic evaluation remains challenging: users experience search results at the page level, while the applicable evaluation criteria are multi-dimensional and continuously evolving. Packing all evaluation criteria into a unified prompt introduces irrelevant context and potential criterion interference, whereas internalizing them through post-training tightly couples rule updates with costly model retraining cycles. To address these issues, we propose Skill-routed Evaluation with Evolvable Knowledge (SEEK). Specifically, SEEK externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result …
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing and behavioral steering, yet they are rarely tested against known failure modes of $L_1$-regularized probing in correlated, high-dimensional feature spaces. We propose a five-step diagnostic protocol covering feature correlation, bootstrap stability, sparse versus dense ranking disagreement, intervention baselines, and cross-dataset evaluation as a minimum standard for sparse-neuron localization claims. We investigate prior work using our proposed approach, specifically on H-neurons using open-source LLMs across TriviaQA, BioASQ, and NQ-Open datasets. Our results demonstr…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening
Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LB…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification
We present TTLab's submission to the AlexandriaX-2026 Subtask~3 on Arabic MT error span detection and classification. Our system frames the task as token-level classification over surface forms, preserving character offsets to ensure exact alignment with the evaluation metric. To handle severe label imbalance, we employ a focal loss with class weighting and dialect-specific decoding thresholds. Among six Arabic pre-trained encoders, MARBERTv2 achieves the best overall performance of 40.8 and 40.91 on the development and test set, respectively, ranking $\nth{3}$ out of all participating teams. While our system localizes error spans effectively, classification of rare error types remains challenging, highlighting the need for data augmentation for tail categories. The code is available at ${\href{https://github.com/ENTAILab/arabic-dialectal-mt-error-span-detection}{\faGithub~ TTLab at Ale…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
StudentBench: AI and human tutoring yield equivalent GRE learning gains
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated l…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better o…
1 matériasInvestimentoBR: Não identificado
Chips e hardwaretrend 34
Nvidia-backed Nscale enterrou acordo com a ByteDance na tentativa de IPO de $35bn - Financial Times
Nvidia-backed Nscale enterrou acordo com a ByteDance na tentativa de IPO de $35bn Financial Times
1 matériasInvestimentoBR: Referência estratégica
Pesquisatrend 34
EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations
Large language models are increasingly used to generate SystemVerilog Assertions from natural-language specifica- tions and register-transfer-level designs. Existing datasets and benchmarks support important goals such as large- scale training, formal evaluation, specification-to-assertion generation, and mutation-based testing. A complemen- tary need is to study whether a generated assertion cap- tures externally observable behavior or depends on inci- dental details of one RTL implementation. We present EquivSVA, a formally verified dataset organized around behavior families. Each family contains four structurally distinct RTL implementations of the same externally ob- servable behavior, shared interface-level gold properties, three controlled mutants, and formal-validation evidence. EquivSVA contains 120 behavior families across 12 cat- egories, 480 reference RTL implementations, 914…
1 matériasInvestimentoBR: Não identificado
Chips e hardwaretrend 34
A NVIDIA está financiando seu próprio rally de ações? - Trefis
A NVIDIA está financiando seu próprio rally de ações? Trefis
1 matériasInvestimentoBR: Referência estratégica
Pesquisatrend 34
Optimal Sequential Annotations for Off-Policy Evaluation
Offline reinforcement learning and off-policy evaluation evaluates dynamic treatment rules based on retrospectively collected data prior to deployment. In recent AI applications, state and reward information is recorded as complex text or image, which recent AI advancements such as LLM-as-a-judge can label with unknown bias. Expert annotation may be available but at a higher cost. For example, safety classification via cheap but imperfect classifiers vs. expensive expert review. We show how a limited budget for ground-truth data-annotation can be used via doubly-robust OPE with missing rewards, and we optimize variance-optimal annotation probabilities for sequential off-policy evaluation, where the target policy value is estimated from annotated data. We characterize the optimal annotation probabilities for sequential forward-monotone annotation protocols, and provide a feasible batch-a…
1 matériasInvestimentoBR: Não identificado
Pesquisatrend 34
Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation
A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some models are accepted and return calls as text, some return native tool_calls, while Phi-3 and Gemma-3 are rejected before inference. In our harness, rejection and retry exhaustion are not preserved as structured failure metadata, so downstream analysis can misclassify them as model non-calls and naively report 0% fidelity. Adding a text tool list while retaining the native channel recovers much of the measured fidelity for accepted models, whereas a uniform text protocol reduces fidelity for Llama…
1 matériasInvestimentoBR: Referência estratégica
Prefere começar pelo PDF de hoje? Baixe o panorama.