Prova formal entra no abismo
AI com Lean em problema do milênio é sinal de fronteira matemática verificável.
fonte: On the Navier–Stokes Millennium Prize Problem
AI com Lean em problema do milênio é sinal de fronteira matemática verificável.
AI com Lean em problema do milênio é sinal de fronteira matemática verificável.
fonte: On the Navier–Stokes Millennium Prize Problem
Evoluir prompts, tools e hooks com correção on-policy pode baratear agents especializados.
fonte: Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Automação de pós-exploração Linux por LLMs toca direto o threat model de agents locais.
fonte: PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation
Coordenação emergente via wiki editável alerta para sandbox, audit log e canais laterais.
fonte: Copying explains the collective behavior of AI agents in the wild
Power monitoring vira métrica verificável de treino frontier, ponte rara entre datacenter e governança.
fonte: FLOP Around and Find Out: LLM Training Workload Size Estimation With Power Monitoring
Grafos deixam planos de agentes auditáveis, mas lembram evolução incremental de workflows.
fonte: Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Foca no gargalo real de coding agents quando testes gerados validam o patch errado.
fonte: ExecCritic: Learn to Test, Test to Improve for Coding Agents
Audita benchmarks cyber de LLMs e mostra que a metodologia altera o placar.
fonte: Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
Truque numérico usa FP8 para multiplicação densa com precisão efetiva maior.
fonte: Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores
Ataca extrapolação além do treino, tema relevante para modelos long-context eficientes.
fonte: Learning Length-Extrapolatable Recurrent Models