Red Hat Ai at SMBtech

Red Hat Adds Agent Operations And Model-As-A-Service To AI Platform With Version 3.4

Surprisingly Useful AI Article Enhancements

Red Hat has announced version 3.4 of its AI platform, introducing tools for managing autonomous AI agents at scale alongside a new Model-as-a-Service capability designed to give developers governed access to curated models through a single interface.

The release, unveiled at Red Hat Summit in Atlanta, is pitched as a “metal-to-agent” platform that spans from hardware infrastructure through to the agentic workflows running on top of it. Red Hat is targeting the gap between AI experimentation and production-grade deployment, a transition that many enterprises are struggling to make.

Red Hat AI 3.4 is expected to be available later this month.

Model-as-a-Service for governed access

The centrepiece of the release is Model-as-a-Service (MaaS), which provides platform engineers with a way to deliver curated, validated models through API endpoints using standard OpenAI-compatible interfaces. Administrators can track consumption and enforce policies, while developers get a consistent way to access models without navigating infrastructure complexity.

The feature supports unified governance of both internal models and external APIs, integrated with identity provider-based authentication.

Underneath, the platform uses the vLLM inference server and llm-d distributed inference engine for model serving. New in this release, request prioritisation allows interactive and background traffic to share the same endpoint, with latency-sensitive requests processed first under load. Speculative decoding support, now generally available, improves response speeds by two to three times with what Red Hat describes as minimal quality impact.

Joe Fernandes, Vice President and General Manager of Red Hat’s AI Business Unit, framed the release around operational control.

“We are defining the open standard for how the enterprise executes AI,” Fernandes explained. “By providing a hardened, metal-to-agent foundation for AI inference, MaaS and AgentOps, Red Hat provides the operational assurance organisations need to innovate at scale while maintaining rigorous control.”

AgentOps for autonomous systems

Red Hat AI 3.4 introduces what it calls AgentOps, a set of tools for managing AI agents from development through to production. The tooling includes integrated tracing, observability, cryptographic identity management and lifecycle management for agents.

The rationale is straightforward: as agents operate with increasing independence, the lack of visibility into their decision-making creates security and governance risks. The platform is designed to trace actions, reasoning steps and tool calls, making it possible to audit how an agent arrived at a particular outcome.

Cryptographic identity management, built on SPIFFE and SPIRE, replaces static hardcoded keys with short-lived tokens. This ties agentic actions to a verified identity and supports least-privilege operations for autonomous agents across the stack.

Safety testing and red-teaming built in

The release integrates automated adversarial scanning directly into the development lifecycle. Using technology from Chatterbox Labs and the open source Garak project, the platform can screen models and agentic systems for risks including jailbreaks, prompt injections and bias.

Nvidia NeMo Guardrails provides run-time safety. An evaluation hub offers a framework-agnostic control plane for benchmarking large language models, AI applications and agents against quality, accuracy and risk metrics.

John Fanelli, Vice President of Enterprise Software at Nvidia, pointed to the governance requirements of autonomous agents.

“Autonomous, long-running agents in the enterprise demand a new level of infrastructure control and security to ensure trustworthy operations at scale,” Fanelli observed. “Red Hat AI Factory with Nvidia provides a unified, open source-driven foundation that gives developers and operators the governance and confidence necessary for the agentic future.”

Prompt management and MLflow integration

Red Hat AI 3.4 introduces prompt management, treating prompts as first-class data assets stored in a central registry. The idea is to give both developers and administrators a single source of truth for the inputs driving models and agents.

The platform integrates MLflow for experiment tracking, artifact management and end-to-end tracing of LLM calls, reasoning steps, tool execution, model responses and token usage via OpenTelemetry. This applies to both generative AI and traditional predictive AI and machine learning use cases.

Automated tools including AutoRAG and AutoML handle tasks such as selecting retrieval strategies for specific datasets and building predictive models.

Extended hardware and cloud support

The release adds day-zero support for Nvidia Blackwell GPUs and AMD MI325X architectures. Red Hat AI Inference now extends beyond Red Hat OpenShift to additional Kubernetes services including CoreWeave and Azure, as well as a new deployment on IBM Cloud.

Urvashi Chowdhary, Vice President of Product Management for AI Services at CoreWeave, described the collaboration as focused on deployment consistency.

“Together, we’ve delivered a deployment blueprint for Red Hat AI Inference on CoreWeave Kubernetes Service to run the same inference stack on-prem and in the cloud, with Kubernetes-native control and production-grade performance,” Chowdhary noted.

Bridging builders and operators

Red Hat is framing the broader challenge as a friction problem between AI developers and infrastructure administrators. The company argues that many organisations recognise the need to move from being “token consumers” to “token providers” in order to manage costs and support private, sovereign AI use cases. But without a unified approach that aligns both roles, infrastructure access barriers slow innovation while ungoverned workarounds introduce risk.

Red Hat AI 3.4 is the company’s answer to that tension, providing a single platform that gives developers self-service access to models and agents while giving operators the governance, tracing and security controls they need to keep things in check.

Last Updated on May 13, 2026 by Nick Ross

Surprisingly Useful AI Article Enhancements

Sign-up to the SMBtech Daily Newsletter

We will not spam you. You can easily unsubscribe any time. Read our privacy policy.