One interface for every model. Authentication, routing, streaming, retries, caching — handled.
Quick Start · Features · Docs · Examples · Providers · Architecture · Contributing
flux is the LLM provider runtime that powers the rho coding agent. It handles everything between your application and LLM APIs — authentication, model resolution, streaming, retries, rate limiting, and caching.
When your app calls a model, flux figures out which provider to use, how to talk to it, and how to stream the response back. Switch from Anthropic to Ollama? flux handles the translation. API returns 529? flux retries with backoff. Response hits max_tokens? flux continues automatically.
Your app never talks to an LLM API directly. flux does.
Rho is the product face: it owns UX, agent orchestration, tools, permissions,
sessions, and product semantics. Flux is the provider engine: it owns
credentials, catalog and route resolution, provider transports, normalized
streams, retry/fallback, usage, and provider telemetry. Rho integrates through
the stable engine facade rather than assembling Flux's internal
provider packages.
flux is a Rho support engine. Keep the dependency edge one-way.
Hosts may import exactly four packages:
| Package | Carries |
|---|---|
engine |
the stable host-facing facade |
llm |
host-facing DTOs and the Provider port that engine re-exports as aliases |
graph |
the portable execution-graph vocabulary |
tools |
tool-call and tool-result contracts |
Everything else is engine-internal: provider, catalog, config,
credentials, router, runtime, and their subpackages are not shared
contracts. Enforced by rho/scripts/check-flux-engine-boundary.sh
and two Go AST tests in rho/internal/testaudit/.
Exception: a fixed set of engine-internal symbols is reachable through the
facade (for example credentials.Store behind engine.Options.SecretStore
and operationsgraph.Input behind engine.OperationsGraphInput). Those
symbols are frozen as part of the contract; see
Frozen engine-internal types.
engine/host_surface_test.go fails when that set changes.
- do not import
rho/internal/* - do not import the removed legacy path
rho/shared/types
go get github.com/GrayCodeAI/fluxRequires Go 1.26+ and a configured provider credential. Direct dependencies:
UUID, tiktoken tokenizer, OS keyring, OpenTelemetry, pure-Go SQLite, and gRPC
(linked only into -tags grpc builds of internal/grpc).
import (
"context"
"fmt"
"github.com/GrayCodeAI/flux/engine"
)
// Hosts (like rho) must use the stable engine facade
eng, err := engine.New(engine.Options{})
if err != nil { panic(err) }
sr, err := eng.Stream(context.Background(), engine.GenerateRequest{
Messages: []engine.Message{{Role: "user", Content: "Hello"}},
Requirements: engine.Requirements{Streaming: true},
Preference: engine.Preference{
PreferredProvider: "anthropic",
PreferredModelID: "anthropic/claude-sonnet-4-6",
},
})
if err != nil { panic(err) }
defer sr.Close()
for sr.Next() {
if evt := sr.Event(); evt.Type == engine.EventContentDelta {
fmt.Print(evt.Content)
}
}
if err := sr.Err(); err != nil { panic(err) }Provider construction lives under provider/; hosts use the
stable engine contract. See
docs/architecture/HOST-ENGINE-BOUNDARY.md
and the feature-oriented architecture.
Automatically detects and routes to the right provider based on environment variables, config files, or explicit selection.
Maps abstract tiers (opus/sonnet/haiku) to concrete model IDs per provider. Ships with an embedded catalog of pricing, context windows, and capabilities.
Parses SSE for Anthropic and OpenAI formats — text, tool calls, and thinking blocks.
- Retries on 429/500/529 with exponential backoff and
Retry-Aftersupport - Auto-continuation when
stop_reason == max_tokens - Provider fallback chains for high availability
Token bucket rate limiter per provider — prevents hitting API limits before they happen.
- Response caching with configurable TTL
- Semantic similarity caching for repeated prompts
- Anthropic prompt caching breakpoints on system prompt and conversation prefix
Built-in cost estimation per call, with per-provider pricing from the embedded model catalog.
Passes reasoning_effort and Anthropic extended-thinking thinking_budget_tokens through to capable models — omitted when unset.
GitHub OIDC keyless authentication for cloud deployments — mints a short-lived token in GitHub Actions and exchanges it for AWS Bedrock (STS AssumeRoleWithWebIdentity) or GCP Vertex (Workload Identity Federation) credentials, no stored secrets.
Serves POST /v1/chat/completions so existing OpenAI SDK clients can talk to flux unchanged.
Named routing strategies beyond weighted random: simple-shuffle, least-busy, latency-based, cost-based, and usage-based.
Distributed CacheBackend interface (in-memory default, RESP/Redis-capable) and an AuditSink interface (no-op default, JSONL file sink) for privacy-preserving call metadata.
Named primary / weak / editor model slots with fallback to primary, plus an LLM summarizing condenser that shrinks long conversation histories using the weak model.
POST /rerank endpoint (provider-backed with lexical fallback) and a GET /ready readiness probe alongside the existing health check.
internal/grpc holds an optional gRPC transport behind the grpc build tag. It serves flux.v1.ChatService/Chat with a registered json content subtype (no .proto files or generated stubs; clients call with grpc.CallContentSubtype("json")), backed by EngineChatService over conversation.Engine. The package is internal, so hosts cannot import it, and nothing in flux starts it. google.golang.org/grpc is a direct requirement in go.mod, so it appears in consumers' module graphs, but only -tags grpc builds link it.
Detailed documentation is available in the docs/ directory:
- Architecture — System design, data flow, and reliability features
- Provider Setup Guide — Credential configuration and provider setup
- Dynamic Model Discovery — Live model discovery architecture
Runnable examples are in the examples/ directory:
- Basic Chat — Simple synchronous chat
- Streaming — SSE streaming with event handling
- Multi-Provider — Fallback chains across providers
Run any example with:
ANTHROPIC_API_KEY=sk-... go run ./examples/basic/28 provider gateways in catalog/registry/providers.go (rho /config uses the same list), listed in registry SortOrder. catalog/registry/docs_test.go fails when this table, the count, or .env.example drift from the registry.
| Provider | ID | Env variable |
|---|---|---|
| Agnes | agnes |
AGNES_API_KEY |
| Amazon Bedrock | bedrock |
AWS_SECRET_ACCESS_KEY (+ AWS_ACCESS_KEY_ID, AWS_SESSION_TOKEN) |
| Anthropic | anthropic |
ANTHROPIC_API_KEY |
| Azure OpenAI | azure |
AZURE_OPENAI_API_KEY (+ AZURE_OPENAI_ENDPOINT) |
| CanopyWave | canopywave |
CANOPYWAVE_API_KEY |
| ClinePass | clinepass |
CLINE_API_KEY |
| Concentrate | concentrate |
CONCENTRATE_API_KEY |
| DeepSeek | deepseek |
DEEPSEEK_API_KEY |
| Google Gemini | gemini |
GEMINI_API_KEY |
| Groq | groq |
GROQ_API_KEY |
| Kimi (Moonshot) | kimi |
MOONSHOT_API_KEY |
| LongCat | longcat |
LONGCAT_API_KEY |
| MiniMax — Pay-as-you-go | minimax_payg |
MINIMAX_PAYG_API_KEY |
| MiniMax — Token Plan | minimax_token_plan |
MINIMAX_TOKEN_PLAN_API_KEY |
| OpenAI | openai |
OPENAI_API_KEY |
| OpenCode Go | opencodego |
OPENCODEGO_API_KEY |
| OpenRouter | openrouter |
OPENROUTER_API_KEY |
| Ollama | ollama |
OLLAMA_BASE_URL (local; no API key) |
| Poolside | poolside |
POOLSIDE_API_KEY |
| Vertex AI | vertex |
VERTEX_ACCESS_TOKEN (or GOOGLE_OAUTH_ACCESS_TOKEN) |
| xAI (Grok) | grok |
XAI_API_KEY |
| Xiaomi (MiMo) Pay-as-you-go | xiaomi_mimo_payg |
XIAOMI_MIMO_PAYG_API_KEY |
| Xiaomi (MiMo) Token Plan | xiaomi_mimo_token_plan |
XIAOMI_MIMO_TOKEN_PLAN_API_KEY (+ region cn / sgp / ams) |
| Z.AI — Coding Plan | zai_coding |
ZAI_CODING_API_KEY (+ region international / cn) |
| Z.AI — Pay-as-you-go | zai_payg |
ZAI_API_KEY (+ region international / cn) |
| StepFun | stepfun |
STEP_API_KEY (+ region global / cn) |
| OpenGateway | opengateway |
OPENGATEWAY_API_KEY |
| Fireworks AI | fireworks |
FIREWORKS_API_KEY |
Runtime auto-detection uses a separate priority order (config.APIProviderDetectionOrder) when no deployment is pinned.
resp, err := c.Chat(ctx, messages, llm.ChatOptions{
Model: "gpt-4o",
})// Auto-continues when max_tokens is hit
resp, err := provider.ChatWithContinuation(ctx, provider, messages,
llm.ChatOptions{Model: model},
core.DefaultContinuationConfig(),
)mock := provider.NewMockProvider(provider.MockModeFixed)
mock.Response = "Here is the code you asked for..."
resp, _ := mock.Chat(ctx, messages, opts)
// No real API calls — perfect for testscat := catalog.DefaultModelCatalog()
// Get the best model for a tier
model := catalog.GetPreferredProviderModel("anthropic", catalog.TierSonnet, &cat)
// → "claude-sonnet-4-6"
// Check deprecation warnings
warn := catalog.GetModelDeprecationWarning("claude-3-7-sonnet", "anthropic")cfg := config.LoadProviderConfig("") // load from disk
config.ApplyProviderConfigToEnv(cfg, false, nil) // apply to environment
config.SaveProviderConfig(cfg, "") // save changesflux/
├── engine/ # Stable host-facing facade (hosts import engine, llm, graph, tools)
├── llm/ # Host-facing DTOs and the Provider port that engine re-exports
├── graph/ # Portable execution-graph vocabulary
├── tools/ # Tool-call and tool-result contracts
├── provider/ # Provider runtime composition root (FluxClient)
│ ├── core/ # Provider-neutral wire, stream, retry, and transport primitives
│ ├── adapters/ # Provider protocol adapters and construction registry
│ ├── resilience/ # Rate limits, continuation, guardrails, and error policy
│ ├── cache/ # Response and semantic caches
│ ├── batch/ # Batch execution
│ ├── embeddings/ # Embedding clients, cache, and defaults
│ ├── media/ # Image and audio clients, structured prompts
│ ├── extraction/ # Structured extraction
│ ├── observability/ # Usage, cost, metrics, tracing, and recording
│ └── testkit/ # Mock provider for tests
├── catalog/ # Model catalog & tier system
│ ├── registry/ # Provider registry (single source of truth for providers)
│ ├── discover/ # Model discovery
│ ├── live/ # Live model listing per provider
│ ├── capabilities/ # Capability and deprecation data
│ └── concentrate/ opencodego/ opengateway/ xiaomi/ zai/ # Gateway-specific helpers
├── config/ # Provider configuration & routing
│ └── credential/ # Credential file management
├── credentials/ # Keyring/env credential stores and OIDC keyless auth
├── router/ # Routing strategies, deployment router, circuit breakers
│ └── controlplane/ # Versioned, signed peer manifests and replicas
├── runtime/ # Engine-internal provider/model/credential resolution
├── setup/ # Catalog-backed deployment wiring
├── operationsgraph/ # Privacy-safe route and generation telemetry projection
├── conversation/ # Conversation engine with branching
├── storage/ # SQLite conversation DAG store, virtual keys, budgets
├── codeagent/ # Code agent retry & fallback strategies
├── verify/ # Provider conformance harness
├── types/ # Shared message types & API errors
├── constants/ # API limits
├── utils/ # Error utilities
├── api/ # OpenAPI spec for internal/api
├── internal/
│ ├── api/ # HTTP API server (library code; no flux binary starts it)
│ ├── cache/ # Cache backends and response cache warmer
│ ├── grpc/ # Optional gRPC transport (build tag grpc)
│ ├── health/ # Provider health checker
│ ├── httputil/ probehttp/ shrink/ # HTTP, probe, and tool-description helpers
│ ├── observability/ # OpenTelemetry spans, metrics, and audit sinks
│ └── sdk/ # Go, Python, TypeScript clients for the internal/api HTTP surface
├── docs/ # Documentation & guides
├── examples/ # Runnable code examples
├── scripts/ # CI guards and helper scripts
└── assets/ # Logo and branding
See docs/ARCHITECTURE.md for detailed system design and data flows.
operationsgraph.Build projects resolved routes and normalized usage into
flux.graph/v1 operations nodes. Provider, model, request ID, and generated
content are represented only by SHA-256 digests; token counts, finish reason,
tool-call count, and deployment-routing state remain queryable.
flux is part of the graycode-eco:
| Component | Repository | Purpose |
|---|---|---|
| rho | GrayCodeAI/rho | AI coding agent |
| flux | This repo | LLM provider runtime |
- Go 1.26+
go build ./... # Verify the library compiles
go test -race ./... # Run all tests with race detector
make ci # Run full CI suite (lint, test, security)
make cover # Generate coverage reportWe welcome contributions! Please see CONTRIBUTING.md for development setup, commit conventions, and the PR process.
Quick start:
- Fork and create a branch:
git checkout -b feat/short-description - Make changes in small, focused commits
- Run
make cilocally - Open a pull request
Use Conventional Commits for commit messages — release-please uses them for versioning.
MIT — see LICENSE for details.
© 2026 GrayCode AI