ai-safety
Published on August 22, 2026 8:42 PM GMT The past few months I've been noticing that several conversations I've had with people working in AI Safety end up on people feeling like the can't talk to anybody outside their field about how the work feels for them. Time and time I'm hearing that the problem is not on if their friends agree or not with them, but rather that they don't have enough contex…
AI Agent Observability, Evaluation, Governance & Policy Enforcement Enforce, not just observe. How many AI agents are running right now? Which agent generated the most LLM cost? Did any agent expose customer data? Did any agent use a restricted LLM model? Is this prompt better than production? Can you prove compliance? $ traccia.init() one init call — trace, cost, and govern every agent One Dashb…

Researchers say the new ‘Cryptographic Context Injection’ technique conceals malicious instructions until they are decrypted inside a trusted execution environment. The post Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini appeared first on SecurityWeek .
Master AI Red Teaming: Learn how to simulate attacks, use PyRIT, and stop prompt injections to secure your Generative AI systems effectively. The post The Comprehensive Guide to AI Red Teaming for Generative Systems appeared first on SourceTrail .
arXiv:2608.19816v1 Announce Type: new Abstract: Decision makers need sufficient understanding to make good decisions about complex AI systems. However, AI deployment decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understandi…

The deal extends Fortinet's reach into AI runtime protection as companies race to secure agents, models and tools across production systems.

Adversa AI has disclosed an attack technique that it says can cause xAI's Grok chatbot to send a user's name, approximate location, subscription tier, and the prompts from the ongoing conversation to an attacker-controlled server after the user asks it to summarize an ordinary web page. The AI security company, which has codenamed the technique "Cryptographic Context Injection," said the

Atalanta's Argo product is now being used to prove the resilience of Viasat’s satellite communications network. The post AI-Assisted Tool Helped Secure Satellite Communication System After 2022 Russian Hacking appeared first on SecurityWeek .

The action taken by OpenAI comes in light of the Hugging Face incident and the discovery of the Astra model’s advanced capabilities. The post OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses appeared first on SecurityWeek .
This practical EU AI Act guide explains whether the Act applies to you, what to assess, and how to approach compliance throughout the project. In 2025, one in five EU enterprises used AI, up from 13.5% in 2024. That growth now carries compliance implications: bans on certain AI practices, AI literacy duties, general-purpose AI model ... Read more
There's a moment every developer hits the first time they connect an AI assistant to a real database: it works beautifully, the model writes a clean SELECT , you get your answer in seconds — and then a small, cold thought arrives. What if it had written DELETE instead? That worry is healthy. An AI agent that can query your production database is also, by default, an AI agent that can UPDATE , DRO…
arXiv:2608.18357v1 Announce Type: new Abstract: Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and officia…
OpenAI is tightening security controls and rethinking its Preparedness Framework as advanced AI models create new cyber risks and computing costs.

Agent Safety is a Box Keep a lid on it. Before we start, let’s cover some terms so we’re thinking about the same thing. This is a post about AI agents, which I’ll define (riffing off Simon Willison 1 ) as: An AI agent runs models and tools in a loop to achieve a goal. Here, goals can include coding, customer service, proving theorems, cloud operations , or many other things. These agents can be i…

The partnership will focus on how AI products respond to self-harm, suicidality and emotional manipulation involving young people Crisis Text Line will help the Youth AI Safety Institute develop standards for severe mental health risks in AI products Crisis Text Line has joined Common Sense Media’s Youth AI Safety Institute as a founding standards partner, bringing evidence from crisis conversati…

research.ioSign up to keep scrolling
Create your feed subscriptions, save articles, keep scrolling.















