Trafy
● Featured today

OpenAI says it slowed Astra model development over security concerns | TechCrunch

OpenAI said this model, which is still in development, reached its "critical cybersecurity threshold," meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems.

StartupsKirsten Korosec·Aug 7, 2026·2 min read
OpenAI says it slowed Astra model development over security concerns | TechCrunchTechCrunch AI

Latest News

After Rippling blew millions on AI in months, it built an employee ROI tool | TechCrunchStartups

After Rippling blew millions on AI in months, it built an employee ROI tool | TechCrunch

After its own AI usage wake-up call, Rippling this week unveiled AI Spend Console, a product that tracks individual and team employee AI spending.

TechCrunch AI · Aug 7, 2026
5 min
Fenix Flexin isn’t even denying using AI to make ‘Rubberz’ anymoreStartups

Fenix Flexin isn’t even denying using AI to make ‘Rubberz’ anymore

Now Fenix’s story is that he ‘never said’ he didn’t use AI to make the song.

The Verge AI · Aug 7, 2026
3 min
Watching Roku’s AI channel is like eating from a troughStartups

Watching Roku’s AI channel is like eating from a trough

Roku City has seen better days.

The Verge AI · Aug 7, 2026
5 min
OpenAI puts the brakes on a new model because it’s supposedly too powerfulStartups

OpenAI puts the brakes on a new model because it’s supposedly too powerful

The new model wasn’t involved in the Hugging Face breach, OpenAI says.

The Verge AI · Aug 7, 2026
2 min
TutorMoments: Do AI tutors know when to help and when to hold back?Open Source

TutorMoments: Do AI tutors know when to help and when to hold back?

A Blog post by Ai2 on Hugging Face

Hugging Face Blog · Aug 7, 2026
8 min
What’s behind the Google AI shake-upStartups

What’s behind the Google AI shake-up

On The Vergecast: The model wars, the scourge of ‘time spent,’ and the ads coming to your BMW.

The Verge AI · Aug 7, 2026
3 min
Cloudflare launches Kitesurf, a browser built for AI agents | TechCrunchStartups

Cloudflare launches Kitesurf, a browser built for AI agents | TechCrunch

Kitesurf is a cloud-hosted browser designed for AI agents instead of people. It uses less computing power than Chromium for common automation tasks, helping developers build browser-based AI agents more efficiently.

TechCrunch AI · Aug 7, 2026
2 min
Airbnb says AI is helping it ship features faster as it tests a new search function | TechCrunchStartups

Airbnb says AI is helping it ship features faster as it tests a new search function | TechCrunch

Airbnb will debut a new AI-powered search experience with a toggle.

TechCrunch AI · Aug 7, 2026
3 min
Jill Lepore on the ‘Artificial State’ and why Silicon Valley's leaders are bad sci-fi readersStartups

Jill Lepore on the ‘Artificial State’ and why Silicon Valley's leaders are bad sci-fi readers

Historian Jill Lepore has a theory about why tech companies often use soaring language to describe their products, almost as if they're forming a new government.

TechCrunch AI · Aug 7, 2026
2 min
New Mexico court orders Meta to pay additional $567M in child safety case | TechCrunchStartups

New Mexico court orders Meta to pay additional $567M in child safety case | TechCrunch

Meta's total fine has racked up to $942 million in this case.

TechCrunch AI · Aug 7, 2026
3 min
Improving Fable 5 SafeguardsLlms

Improving Fable 5 Safeguards

We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces fallbacks.

Anthropic News · Aug 7, 2026
7 min
OpenAI's new AI smart speaker will reportedly sell for between $300 and $400 | TechCrunchStartups

OpenAI's new AI smart speaker will reportedly sell for between $300 and $400 | TechCrunch

Additional details about OpenAI's mysterious new AI device make it sound like a pricey smart speaker.

TechCrunch AI · Aug 6, 2026
2 min
Jony Ive’s first OpenAI gadget is reportedly a hockey puck-sized smart speakerStartups

Jony Ive’s first OpenAI gadget is reportedly a hockey puck-sized smart speaker

OpenAI’s first AI device could launch next year.

The Verge AI · Aug 6, 2026
3 min
Learning When to Trust via Selective Context Preference OptimizationResearch

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misle

arXiv (cs.AI) · Aug 6, 2026
4 min
Learning When to Trust via Selective Context Preference OptimizationResearch

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, tr...

Papers with Code · Aug 6, 2026
1 min
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature EngineeringResearch

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubr

arXiv (cs.AI) · Aug 6, 2026
4 min
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of NigeriaResearch

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applications affects digital sovereignty, examined through platform transparency as a key indicator of user awareness and control. Using an interpretive approach, the research combines the forensic analysis of selected Android applications with contextual document analysis to identify AI features

arXiv (cs.AI) · Aug 6, 2026
3 min
An Optimal Agnostic PAC AlgorithmResearch

An Optimal Agnostic PAC Algorithm

Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size $n$, for every $0<δ\le 1/2$, with probability at least $1-δ$, \[ L(\widehat h) \le L^*+ 7\cdot10^8\left( \sqrt{\frac{L^*(d+\log(1/δ))}{n}} +\frac{d+\log(1/δ)}{n} \right). \] This settles the sample complexity of agnostic PAC learning up to universal constants at every fixed $L^*$, matching the lower bounds of Devroye, Györfi, and Lugosi [A Probabilistic

arXiv (cs.AI) · Aug 6, 2026
3 min
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information GamesResearch

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-informat

arXiv (cs.AI) · Aug 6, 2026
4 min
The Low Frequency Trap: Video Language Models Fail at Simple Event BookkeepingResearch

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and freque

arXiv (cs.AI) · Aug 6, 2026
4 min
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI AgentsResearch

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism seeks to establish the Safe AI paradigm that compute is an effective governance lever. We situate our work as a compliance or commons overlay on a deployer. One governance period is an extensive form game in which verified human stakeholders arrive sequentially and contribute, on a provision

arXiv (cs.AI) · Aug 6, 2026
4 min
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal TasksResearch

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated st

arXiv (cs.LG) · Aug 6, 2026
4 min
Challenges in Evaluating Explanation Methods for Static and Evolving DataResearch

Challenges in Evaluating Explanation Methods for Static and Evolving Data

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift. Experiences with adapting counterfactuals for this problem are discussed. Finally it is related to the challenges of tracking the co-evolution

arXiv (cs.AI) · Aug 6, 2026
3 min
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent TrajectoriesResearch

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with differe

arXiv (cs.AI) · Aug 6, 2026
4 min