AI NEWS DESK

Agents
AI news.

Current agents releases and research, summarized from live sources and linked to the original publisher.

AWS Machine Learning

Introducing Kimi K3 on Amazon Bedrock

Open-weight models are changing the economics of building and deploying AI at scale. Rapid gains in intelligence and efficiency mean companies can match each workload with the right balance of capability, speed, and cost. AWS is building for a future in which organizations can adopt open-weight innovation with the reliability and security required for production. Today, Kimi K3 from Moonshot AI is available on Amazon Bedrock, giving you a powerful new option for coding and knowledge work. According to Moonshot AI,

READ BRIEFING ↗
AWS Machine Learning

Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

Organizations building multi-model agentic AI applications face growing infrastructure complexity. Managing container orchestration, scaling policies, identity, and observability for multiple model types adds operational overhead. Teams often spend more time on infrastructure than on agent logic development. Developers running agentic frameworks on self-managed infrastructure such as Amazon Elastic Container Service (Amazon ECS) with AWS Fargate have full control over their deployment configuration. As agentic work

READ BRIEFING ↗
AWS Machine Learning

The new AgentCore runtime: Elastic, optimized, and consistently fast starts

Agents are no longer experiments. They process claims, write and review code, coordinate across systems, and run for hours without supervision. As agents take on more complex, longer-running work, the infrastructure underneath them must evolve just as fast. We built Amazon Bedrock AgentCore to help developers build, connect, and optimize agents securely at scale. AgentCore runtime, a capability of Amazon Bedrock AgentCore, is the managed compute layer that gives developers a fully managed environment to deploy and

READ BRIEFING ↗
AWS Machine Learning

Deploy Hugging Face models on Amazon SageMaker AI with coding agents

Deploying a Hugging Face model to production means making a dozen decisions: choosing the right serving container for the model’s architecture, confirming the current image tag for your AWS Region, and matching an instance type to the model’s memory footprint. Beyond infrastructure, you must wire autoscaling so you don’t burn GPU hours on an idle endpoint. You also set Amazon CloudWatch alarms that catch silent failures before your users do. After you’ve made those decisions, Amazon SageMaker AI collapses that work

READ BRIEFING ↗
AWS Machine Learning

Selecting a vector store for Amazon Bedrock Knowledge Bases

When building a Retrieval Augmented Generation (RAG) solution with Amazon Bedrock Knowledge Bases , selecting the right vector store impacts performance and cost. Amazon Bedrock Knowledge Bases offers a fully managed option and a customer-managed option where you choose your own vector store. This post focuses on the customer-managed path, comparing the three supported backends: Amazon OpenSearch Service , Amazon Aurora PostgreSQL with pgvector , and Amazon S3 Vectors , a capability of Amazon Simple Storage Service

READ BRIEFING ↗
AWS Machine Learning

A serverless, data-driven Git metrics dashboard using Amazon Quick Sight

Git activity is one of the richest signals engineering teams produce that can provide continuous observability into development analytics. The challenge is extracting these Git metrics at scale, which has traditionally required hand-rolled extract, transform, and load (ETL) jobs, dedicated infrastructure, and ongoing maintenance. Further, with modern development tools becoming more prevalent, teams need a clear way to measure whether these tools are making developers faster, or if the investment is not paying off.

READ BRIEFING ↗
AWS Machine Learning

A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore

Building a working agentic prototype takes an afternoon. Getting it to production is where the work explodes. The moment an agent has to serve more than one user, a new layer of engineering appears and it’s critical to tell whether the agent is doing the right thing on real traffic. Concurrency, session isolation, identity, persistent state, scaling, and guardrails are all layers that most teams rebuild from scratch every time, even though a standardized platform for each agent is a more repeatable approach. The re

READ BRIEFING ↗
AWS Machine Learning

How MRH Trowe enabled secure self-service AI agents in financial services

Basic AI chat isn’t enough for financial services organizations that need secure, self-service AI agents. In financial services, employees need AI that can work with internal systems and sensitive client data, stay inside a governed environment, and remain auditable and cost-transparent. All of this must happen without every team standing up its own tools. This post shows how MRH Trowe , one of Germany’s leading commercial and industrial insurance brokers, gave approximately 400 employees secure access to self-serv

READ BRIEFING ↗
NVIDIA Developer

How to Use AI Agents to Prepare 3D Scenes for Simulation

Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in...

READ BRIEFING ↗
NVIDIA Developer

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through...

READ BRIEFING ↗
AWS Machine Learning

Improving HCLS AI reasoning with open-source agent skills

AI agents built on foundation models (FMs) often misapply healthcare and life sciences (HCLS) decision frameworks, even when they’ve seen the guidelines in training and in the system prompt. Ask an agent to classify a TP53 missense variant using ACMG/AMP criteria. It will cite the correct framework but misapply evidence categories, skip population frequency thresholds, or hallucinate computational predictor scores. The model knows facts but lacks the structured reasoning procedures that domain practitioners interna

READ BRIEFING ↗
AWS Machine Learning

Fault tolerant distributed training on Amazon EKS using NVRx

Large-scale distributed training jobs run for hours or days across dozens of nodes. At that scale and duration, interruptions are statistically inevitable: network partitions, memory errors, software exceptions, or infrastructure events will eventually disrupt at least one worker. A single GPU fault triggers a cascade: NVIDIA Collective Communication Library (NCCL) timeouts propagate to healthy workers, pods crash and restart out of sync, and your cluster burns expensive GPU hours while making zero training progres

READ BRIEFING ↗
NVIDIA Developer

Translating CUDA Tile Operations from Python to Rust Using Agentic AI

cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to...

READ BRIEFING ↗
AWS Machine Learning

Optimizing agent system prompts with Amazon Bedrock AgentCore

In a previous launch post , we introduced AgentCore optimization, a capability of Amazon Bedrock AgentCore that can help you improve the quality of your agents. Improving a low-scoring agent has traditionally been a manual process. You review long traces to find where the agent goes wrong, tune individual components such as prompts, tool descriptions, and skills, and rerun evaluations to check for improvement. With AgentCore optimization, you can use production traces to propose configuration changes, validate them

READ BRIEFING ↗
AWS Machine Learning

Build a serverless PII redaction pipeline with Amazon Bedrock Data Automation

Organizations that process thousands of scanned documents daily, including medical forms, insurance claims, and financial records, face a recurring compliance need: personally identifiable information (PII) redaction before documents are shared with third parties or processed downstream. Manual redaction doesn’t scale: It consumes staff hours, introduces human error, and creates compliance exposure. Redaction is also a precision problem, in addition to a detection problem. A single page can contain multiple names,

READ BRIEFING ↗
OpenAI

Reimagining advertising with AI

Explore new AI-powered advertising experiences from OpenAI, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify.

READ BRIEFING ↗
OpenAI

Hex turns complex analysis into visual reports with GPT‑6 Astra

GPT-6 Astra helps Hex’s data agents turn answers into interactive visualizations that employees are proud to share.

READ BRIEFING ↗
Hugging Face

Your Agent Aced the Task. Will It Do It Again?

Read Hugging Face’s complete announcement and technical details at the original link.

READ BRIEFING ↗
OpenAI

Now everyone can put data to work

Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.

READ BRIEFING ↗
NVIDIA Developer

Building a Memory-Driven Agent with NVIDIA NemoClaw

Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it...

READ BRIEFING ↗
NVIDIA Developer

Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson

Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...

READ BRIEFING ↗
Hugging Face

Give Your Coding Agents a Memory You Own

Read Hugging Face’s complete announcement and technical details at the original link.

READ BRIEFING ↗
Hugging Face

Training a coding model to paint watercolours with TRL and OpenEnv

Read Hugging Face’s complete announcement and technical details at the original link.

READ BRIEFING ↗
Google Research

AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

Human-Computer Interaction and Visualization

READ BRIEFING ↗
Microsoft Research

Orchard: An open framework for scalable agentic AI

At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains. The same Orchard infrastructure supports software-engineering, web-navigation, and personal-assistant agents, and can train them directly inside real deployment harnesses such as Codex, OpenClaw, and ZeroClaw—letting researchers reuse environments, data pipelines, and evaluation workflows across tasks. Or

READ BRIEFING ↗
Microsoft Research

Echoverse: Deep, evolving environments for computer-use agents

Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train. At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application’s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Train

READ BRIEFING ↗
Google Research

SymptomAI: Towards a conversational AI agent for everyday symptom assessment

General Science

READ BRIEFING ↗
Microsoft Research

Verifying Rust cryptography in SymCrypt, from standards to code

How Rust, Lean, Aeneas, and AI agents are helping scale formal verification for production cryptographic algorithms At a glance SymCrypt develops new verified cryptography using Rust, Aeneas, and Lean to provide higher security assurance. We prove that their code safely and correctly implements standard algorithms, notably for post-quantum cryptography. We are releasing verified code, specs, properties, and proofs initially for SHA-3 and ML-KEM. Aeneas allows verifying a large subset of Rust code and provides effic

READ BRIEFING ↗
Microsoft Research

Flint: A visualization language for the AI era

At a glance Polished charts from simple specs . Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications. Semantic types guide design . Flint leverages semantic data types to express meanings of data. They help the compiler choose appropriate scales, baselines, formatting, and color schemes. Layouts adapt to the data . Flint automatically manages sizing, spacing, labels, and layout so charts remain readable as cardinality and density change, without

READ BRIEFING ↗