diff --git a/docs/about-nemo-relay/overview.mdx b/docs/about-nemo-relay/overview.mdx index 85f644d88..d29eeff90 100644 --- a/docs/about-nemo-relay/overview.mdx +++ b/docs/about-nemo-relay/overview.mdx @@ -13,11 +13,11 @@ without rewriting the agent stack you already have. It gives coding agents, applications, framework integrations, middleware, and observability backends a shared runtime for scopes, policy, plugins, and lifecycle events. -Agent systems usually cross several boundaries in one request: an entrypoint +Agent systems usually involve several components in one request: an entry point starts work, a model is called, tools run, subagents can branch off, and observability or policy systems need to understand what happened. Relay gives -those boundaries one runtime contract instead of asking each layer to invent its -own wrappers, trace vocabulary, and cleanup rules. +those components one runtime contract instead of asking each layer to invent +its own wrappers, trace vocabulary, and cleanup rules. ## Integrating With Relay @@ -27,12 +27,13 @@ run, or a framework-specific lifecycle unit. Relay does not replace your agent framework, model provider, application logic, -observability backend, or guardrail authoring system. It gives those systems a -common runtime boundary to meet at. +observability backend, or guardrail authoring system. It connects those systems +through shared scopes, middleware, plugins, and lifecycle events. For how Relay complements OpenTelemetry GenAI conventions and observability or -evaluation products, see [How NeMo Relay Relates to Other Tooling](/about-nemo-relay/ecosystem#how-nemo-relay-relates-to-other-tooling). +evaluation products, refer to +[How NeMo Relay Relates to Other Tooling](/about-nemo-relay/ecosystem#how-nemo-relay-relates-to-other-tooling). The first design question is simple: where can Relay observe or control the real work? The answer determines whether you should use a CLI sidecar, direct SDK @@ -45,75 +46,73 @@ Pick the row closest to what you are trying to do. | Goal | Start With | Why | |---|---|---| | Observe Codex, Claude Code, or Hermes locally | [NeMo Relay CLI](/nemo-relay-cli/about) and [Basic Usage](/nemo-relay-cli/basic-usage) | Relay runs as a local sidecar, forwards hooks, routes provider traffic when configured, and writes observability artifacts without changing application code. | -| Run the smallest binding-specific example | [Quick Start](/getting-started/quick-start) | Use this when you want a minimal Rust, Python, or Node.js workflow before adding Relay to real application code. | -| Instrument application-owned LLM or tool calls | [Instrument Applications](/instrument-applications/about) | Direct SDK instrumentation gives Relay full managed-call semantics around callbacks your code owns. | -| Use LangChain, LangGraph, Deep Agents, or OpenClaw | [Supported Integrations](/supported-integrations/about) | Maintained integrations use public framework or plugin APIs where they preserve enough lifecycle fidelity. | -| Build a framework, host, or provider integration | [Integrate into Frameworks](/integrate-into-frameworks/about) | Integration guidance helps you choose managed wrappers, explicit lifecycle APIs, hook replay, provider codecs, or upstream support. | -| Package reusable exporters, middleware, or policy | [Build Plugins](/build-plugins/about) and [Configure Plugins](/configure-plugins/about) | Plugins are the configuration-driven path for behavior that should be shared across applications or teams. | -| Develop or validate the repository itself | [Development Setup](/contribute/development-setup) and [Testing and Docs](/contribute/testing-and-docs) | Use the contributor workflow when you are changing Relay source, docs, examples, bindings, or integrations. | +| Instrument application-owned LLM or tool calls | [Instrument Applications](/instrument-applications/about) | Direct SDK instrumentation lets Relay run the complete lifecycle and middleware sequence around callbacks your code owns. | +| Use LangChain, LangGraph, Deep Agents, or OpenClaw | [Supported Integrations](/supported-integrations/about) | Maintained integrations use public framework or plugin APIs to capture supported lifecycle events. | +| Configure traces, trajectories, or raw event export | [Observability](/configure-plugins/observability/about) | Exporters consume the same lifecycle event stream and write ATOF, ATIF, OpenTelemetry, or OpenInference output. | -If you are unsure how much Relay you need, capture one boundary first. Confirm -that Relay emits raw lifecycle events, then add normalized exports, middleware, -guardrails, or adaptive behavior. +To evaluate a language binding with the smallest complete example, start with +[Quick Start](/getting-started/quick-start). -## Validate Raw Capture First +Relay records canonical lifecycle events in Agent Trajectory Observability +Format (ATOF). Exporters can write those events directly as +[ATOF JSONL](/configure-plugins/observability/atof), project completed runs into +[Agent Trajectory Interchange Format (ATIF)](/configure-plugins/observability/atif) +trajectories, or translate them into typed OpenTelemetry output. -Start with [Agent Trajectory Observability Format (ATOF) JSONL](/configure-plugins/observability/atof), -the raw canonical event stream. It shows the lifecycle events Relay actually -captured before anything is translated into -[Agent Trajectory Interchange Format (ATIF)](/configure-plugins/observability/atif), -or a typed OpenTelemetry projection. +## Build on Relay -A good first integration process workflow is as follows: +Use these paths when you need to extend Relay instead of using an existing +feature or integration. -1. Create or identify one scope boundary. -2. Capture one LLM, tool, session, or turn boundary. -3. Export ATOF JSONL and inspect the raw event stream. -4. Add ATIF or a typed OpenTelemetry projection when the raw events are trustworthy. -5. Add middleware only when Relay must block, sanitize, rewrite, route, or - replace real execution. +| Goal | Start With | +|---|---| +| Build a framework, host, or provider integration | [Integrate into Frameworks](/integrate-into-frameworks/about) | +| Package reusable exporters, middleware, or policy | [Build Plugins](/build-plugins/about) and [Configure Plugins](/configure-plugins/about) | +| Develop or validate the repository | [Development Setup](/contribute/development-setup) and [Testing and Docs](/contribute/testing-and-docs) | + +Rust is the source of truth for runtime behavior. The Python and Node.js +bindings expose the same core model for primary application use. Go and raw C +FFI are experimental and source-first. + +### Choose How Relay Connects + +Identify where the actual LLM or tool function is invoked. If that invocation +can be routed through NeMo Relay, use managed execution: NeMo Relay runs the +applicable middleware and then invokes the real function. If the framework +retains control but provides before-and-after notifications, translate those +lifecycle hooks into NeMo Relay events. NeMo Relay can then observe the call's +lifecycle but does not execute or control it. If provider-native requests and +responses must be intercepted, route the real provider traffic through the +NeMo Relay gateway and treat NeMo Relay as a production dependency. + + +Do not add behavior to one primary binding without checking Rust, Python, and +Node.js parity. Public behavior should stay consistent across the supported +bindings. + ## Key Features -NeMo Relay offers the following features when you use it with your agent stacks: +NeMo Relay offers the following features for agent applications: +- **Events and subscribers** so ATOF events, ATIF trajectories, and typed + OpenTelemetry output come from the same runtime activity. - **Scopes** so runs, turns, tools, LLM calls, and subagents have clear - ownership, parent-child lineage, cleanup boundaries, and request isolation. + parent-child relationships, automatic cleanup, and request isolation. +- **Marks** so point-in-time events, such as session starts, compaction, or skill + loads, do not require a start and end pair. - **Managed LLM and tool calls** so the same lifecycle and middleware rules apply around each callback. - **Middleware** for the places where Relay must block, sanitize, transform, route, retry, or replace execution. - **Plugins** so reusable observability, guardrail, adaptive, and exporter behavior can be turned on from configuration. -- **Events and subscribers** so raw ATOF, normalized ATIF, and typed - OpenTelemetry output all come from the same runtime stream. Use [Concepts](/about-nemo-relay/concepts) when you want the deeper model for scopes, events, middleware, subscribers, and plugins. -## Developer Background - -Rust is the source of truth for runtime behavior. The Python and Node.js -bindings expose the same core model for primary application use. Go and raw C -FFI are experimental and source-first surfaces. - -First, identify where the actual LLM or tool function is invoked. If that -invocation can be routed through NeMo Relay, use managed execution: NeMo Relay -runs the applicable middleware and then invokes the real function. If the -framework retains control but provides before-and-after notifications, translate -those lifecycle hooks into NeMo Relay events. NeMo Relay can then observe the -call's lifecycle but does not execute or control it. If provider-native requests -and responses must be intercepted, route the real provider traffic through the -NeMo Relay gateway and treat NeMo Relay as a production dependency. - - -Do not add behavior to one primary binding without checking Rust, Python, and -Node.js parity. Public behavior should stay consistent across the supported -runtime surfaces. - - ## Documentation Use the tasks below to build your understanding and set up Relay: @@ -121,66 +120,52 @@ Use the tasks below to build your understanding and set up Relay: | Task | Start With | |---|---| | Install packages | [Installation](/getting-started/installation) | -| Understand the mental model | [Agent Runtime Primer](/about-nemo-relay/agent-runtime-primer) | +| Understand the runtime model | [Agent Runtime Primer](/about-nemo-relay/agent-runtime-primer) | | Configure plugin files | [Plugin Configuration Files](/configure-plugins/plugin-configuration-files) | | Export traces or trajectories | [Observability](/configure-plugins/observability/about) | | Tune performance with adaptive behavior | [Adaptive](/configure-plugins/adaptive/about) | | Debug trace incidents | [Trace Incident Runbook](/resources/troubleshooting/trace-incident-runbook) | | Look up symbols | [APIs](/reference/api) | -## Conceptual Diagram +## How Relay Connects to Your Stack -The diagram below shows how applications, runtime components, and exporters -relate to each other. Scopes define where work belongs, middleware registries -define what runs around that work, and subscribers consume the lifecycle events -that the core emits. +The diagram shows the external flow from application work through Relay to +observability output. Applications and frameworks keep ownership of their real +payloads and callbacks. Relay records lifecycle events for subscribers and +exporters to consume. ```mermaid -flowchart TB - Plugin[Plugin] - App[Application Code / Agent Harness / Agent Framework] - Framework[Framework Integration] - - subgraph Runtime[Runtime] - PluginSystem[Plugin System] - Bindings[Language Bindings] - Core[Rust Core Runtime] - Events[Lifecycle Events] - - subgraph RuntimeState[Runtime State] - Registry[Middleware Registries
what runs around work] - Scope[Scope Stack
where work belongs] - end - - Subs[Subscribers / Exporters] - - PluginSystem --->|installs| Registry - PluginSystem ----->|installs| Subs - Bindings --> Core - Core -->|emits| Events -->|consumed by| Subs - Core -->|updates| Scope - Core -->|resolves| Registry +flowchart LR + Sources[Application Code / Agent Harness / Agent Framework] + Runtime[Relay Runtime] + Consumers[Subscribers / Exporters] + + subgraph Outputs[Outputs and Destinations] + ATOF[ATOF JSONL] + ATIF[ATIF Trajectories] + OTEL[OpenTelemetry] + OI[OpenInference] end - App -->|registers| Plugin - App -->|uses| Framework - App -->|configures/initializes| PluginSystem - App -->|uses| Bindings - Framework -->|calls| Bindings - Plugin -->|registers with| PluginSystem - - class Runtime grey-lightest; - class RuntimeState grey-lightest; - class App purple-lightest; - class Framework yellow-lightest; - class Plugin blue-lightest; - class Bindings green-lightest; - class PluginSystem green-light; - class Core green-light; - class Scope green-light; - class Registry green-light; - class Events green-light; - class Subs green-light; + Sources -->|managed calls, hooks, or routed traffic| Runtime + Runtime -->|canonical lifecycle events| Consumers + Consumers --> ATOF + Consumers --> ATIF + Consumers --> OTEL + Consumers --> OI + + class Sources purple-lightest; + class Runtime green-light; + class Consumers green-lightest; + class Outputs grey-lightest; + class ATOF,ATIF,OTEL,OI blue-lightest; ``` + +Configure the output that matches your destination: + +- [ATOF JSONL](/configure-plugins/observability/atof) for the canonical event stream. +- [ATIF trajectories](/configure-plugins/observability/atif) for replay, analysis, and evaluation. +- [OpenTelemetry](/configure-plugins/observability/opentelemetry) for OTLP-compatible backends. +- [OpenInference](/configure-plugins/observability/openinference) for OpenInference-compatible tracing.