ai-safety

I Built a Capability-Based Security Layer for AI Agents — Here's Why It Matters The Problem Nobody's Talking About AI agents are everywhere now. They book flights, send emails, process payments, and access your codebase. But here's the question nobody asks: Who authorizes which agent can do what? Most people use API keys. An API key is binary — you have it or you don't. If your finance agent's ke…

Published on August 22, 2026 8:42 PM GMT The past few months I've been noticing that several conversations I've had with people working in AI Safety end up on people feeling like the can't talk to anybody outside their field about how the work feels for them. Time and time I'm hearing that the problem is not on if their friends agree or not with them, but rather that they don't have enough contex…

AI Agent Observability, Evaluation, Governance & Policy Enforcement Enforce, not just observe. How many AI agents are running right now? Which agent generated the most LLM cost? Did any agent expose customer data? Did any agent use a restricted LLM model? Is this prompt better than production? Can you prove compliance? $ traccia.init() one init call — trace, cost, and govern every agent One Dashb…

arXiv:2608.19816v1 Announce Type: new Abstract: Decision makers need sufficient understanding to make good decisions about complex AI systems. However, AI deployment decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understandi…

Adversa AI has disclosed an attack technique that it says can cause xAI's Grok chatbot to send a user's name, approximate location, subscription tier, and the prompts from the ongoing conversation to an attacker-controlled server after the user asks it to summarize an ordinary web page. The AI security company, which has codenamed the technique "Cryptographic Context Injection," said the

There's a moment every developer hits the first time they connect an AI assistant to a real database: it works beautifully, the model writes a clean SELECT , you get your answer in seconds — and then a small, cold thought arrives. What if it had written DELETE instead? That worry is healthy. An AI agent that can query your production database is also, by default, an AI agent that can UPDATE , DRO…

arXiv:2608.18357v1 Announce Type: new Abstract: Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and officia…

Agent Safety is a Box Keep a lid on it. Before we start, let’s cover some terms so we’re thinking about the same thing. This is a post about AI agents, which I’ll define (riffing off Simon Willison 1 ) as: An AI agent runs models and tools in a loop to achieve a goal. Here, goals can include coding, customer service, proving theorems, cloud operations , or many other things. These agents can be i…

The partnership will focus on how AI products respond to self-harm, suicidality and emotional manipulation involving young people Crisis Text Line will help the Youth AI Safety Institute develop standards for severe mental health risks in AI products Crisis Text Line has joined Common Sense Media’s Youth AI Safety Institute as a founding standards partner, bringing evidence from crisis conversati…

research.ioresearch.io

Sign up to keep scrolling

Create your feed subscriptions, save articles, keep scrolling.

Already have an account?