Skip to content

Latest commit

 

History

301 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

flux

Universal LLM Provider Runtime

One interface for every model. Authentication, routing, streaming, retries, caching — handled.

Go License CI Release GoDoc

Quick Start · Features · Docs · Examples · Providers · Architecture · Contributing


What is flux

flux is the LLM provider runtime that powers the rho coding agent. It handles everything between your application and LLM APIs — authentication, model resolution, streaming, retries, rate limiting, and caching.

When your app calls a model, flux figures out which provider to use, how to talk to it, and how to stream the response back. Switch from Anthropic to Ollama? flux handles the translation. API returns 529? flux retries with backoff. Response hits max_tokens? flux continues automatically.

Your app never talks to an LLM API directly. flux does.

Rho is the product face: it owns UX, agent orchestration, tools, permissions, sessions, and product semantics. Flux is the provider engine: it owns credentials, catalog and route resolution, provider transports, normalized streams, retry/fallback, usage, and provider telemetry. Rho integrates through the stable engine facade rather than assembling Flux's internal provider packages.

Ecosystem Boundaries

flux is a Rho support engine. Keep the dependency edge one-way.

Hosts may import exactly four packages:

Package Carries
engine the stable host-facing facade
llm host-facing DTOs and the Provider port that engine re-exports as aliases
graph the portable execution-graph vocabulary
tools tool-call and tool-result contracts

Everything else is engine-internal: provider, catalog, config, credentials, router, runtime, and their subpackages are not shared contracts. Enforced by rho/scripts/check-flux-engine-boundary.sh and two Go AST tests in rho/internal/testaudit/.

Exception: a fixed set of engine-internal symbols is reachable through the facade (for example credentials.Store behind engine.Options.SecretStore and operationsgraph.Input behind engine.OperationsGraphInput). Those symbols are frozen as part of the contract; see Frozen engine-internal types. engine/host_surface_test.go fails when that set changes.

  • do not import rho/internal/*
  • do not import the removed legacy path rho/shared/types

Quick Start

go get github.com/GrayCodeAI/flux

Requires Go 1.26+ and a configured provider credential. Direct dependencies: UUID, tiktoken tokenizer, OS keyring, OpenTelemetry, pure-Go SQLite, and gRPC (linked only into -tags grpc builds of internal/grpc).

import (
    "context"
    "fmt"

    "github.com/GrayCodeAI/flux/engine"
)

// Hosts (like rho) must use the stable engine facade
eng, err := engine.New(engine.Options{})
if err != nil { panic(err) }

sr, err := eng.Stream(context.Background(), engine.GenerateRequest{
    Messages: []engine.Message{{Role: "user", Content: "Hello"}},
    Requirements: engine.Requirements{Streaming: true},
    Preference: engine.Preference{
        PreferredProvider: "anthropic",
        PreferredModelID: "anthropic/claude-sonnet-4-6",
    },
})
if err != nil { panic(err) }
defer sr.Close()

for sr.Next() {
    if evt := sr.Event(); evt.Type == engine.EventContentDelta {
        fmt.Print(evt.Content)
    }
}
if err := sr.Err(); err != nil { panic(err) }

Provider construction lives under provider/; hosts use the stable engine contract. See docs/architecture/HOST-ENGINE-BOUNDARY.md and the feature-oriented architecture.

Features

Provider Routing

Automatically detects and routes to the right provider based on environment variables, config files, or explicit selection.

Model Resolution

Maps abstract tiers (opus/sonnet/haiku) to concrete model IDs per provider. Ships with an embedded catalog of pricing, context windows, and capabilities.

Streaming

Parses SSE for Anthropic and OpenAI formats — text, tool calls, and thinking blocks.

Reliability

  • Retries on 429/500/529 with exponential backoff and Retry-After support
  • Auto-continuation when stop_reason == max_tokens
  • Provider fallback chains for high availability

Rate Limiting

Token bucket rate limiter per provider — prevents hitting API limits before they happen.

Caching

  • Response caching with configurable TTL
  • Semantic similarity caching for repeated prompts
  • Anthropic prompt caching breakpoints on system prompt and conversation prefix

Cost Tracking

Built-in cost estimation per call, with per-provider pricing from the embedded model catalog.

Reasoning Controls

Passes reasoning_effort and Anthropic extended-thinking thinking_budget_tokens through to capable models — omitted when unset.

Keyless CI Auth

GitHub OIDC keyless authentication for cloud deployments — mints a short-lived token in GitHub Actions and exchanges it for AWS Bedrock (STS AssumeRoleWithWebIdentity) or GCP Vertex (Workload Identity Federation) credentials, no stored secrets.

OpenAI-Compatible Proxy

Serves POST /v1/chat/completions so existing OpenAI SDK clients can talk to flux unchanged.

Load-Balancing Strategies

Named routing strategies beyond weighted random: simple-shuffle, least-busy, latency-based, cost-based, and usage-based.

Pluggable Cache & Audit Sinks

Distributed CacheBackend interface (in-memory default, RESP/Redis-capable) and an AuditSink interface (no-op default, JSONL file sink) for privacy-preserving call metadata.

Model Role Slots

Named primary / weak / editor model slots with fallback to primary, plus an LLM summarizing condenser that shrinks long conversation histories using the weak model.

Rerank & Readiness

POST /rerank endpoint (provider-backed with lexical fallback) and a GET /ready readiness probe alongside the existing health check.

gRPC Transport (opt-in, internal)

internal/grpc holds an optional gRPC transport behind the grpc build tag. It serves flux.v1.ChatService/Chat with a registered json content subtype (no .proto files or generated stubs; clients call with grpc.CallContentSubtype("json")), backed by EngineChatService over conversation.Engine. The package is internal, so hosts cannot import it, and nothing in flux starts it. google.golang.org/grpc is a direct requirement in go.mod, so it appears in consumers' module graphs, but only -tags grpc builds link it.

Documentation

Detailed documentation is available in the docs/ directory:

Examples

Runnable examples are in the examples/ directory:

Run any example with:

ANTHROPIC_API_KEY=sk-... go run ./examples/basic/

Supported Providers

28 provider gateways in catalog/registry/providers.go (rho /config uses the same list), listed in registry SortOrder. catalog/registry/docs_test.go fails when this table, the count, or .env.example drift from the registry.

Provider ID Env variable
Agnes agnes AGNES_API_KEY
Amazon Bedrock bedrock AWS_SECRET_ACCESS_KEY (+ AWS_ACCESS_KEY_ID, AWS_SESSION_TOKEN)
Anthropic anthropic ANTHROPIC_API_KEY
Azure OpenAI azure AZURE_OPENAI_API_KEY (+ AZURE_OPENAI_ENDPOINT)
CanopyWave canopywave CANOPYWAVE_API_KEY
ClinePass clinepass CLINE_API_KEY
Concentrate concentrate CONCENTRATE_API_KEY
DeepSeek deepseek DEEPSEEK_API_KEY
Google Gemini gemini GEMINI_API_KEY
Groq groq GROQ_API_KEY
Kimi (Moonshot) kimi MOONSHOT_API_KEY
LongCat longcat LONGCAT_API_KEY
MiniMax — Pay-as-you-go minimax_payg MINIMAX_PAYG_API_KEY
MiniMax — Token Plan minimax_token_plan MINIMAX_TOKEN_PLAN_API_KEY
OpenAI openai OPENAI_API_KEY
OpenCode Go opencodego OPENCODEGO_API_KEY
OpenRouter openrouter OPENROUTER_API_KEY
Ollama ollama OLLAMA_BASE_URL (local; no API key)
Poolside poolside POOLSIDE_API_KEY
Vertex AI vertex VERTEX_ACCESS_TOKEN (or GOOGLE_OAUTH_ACCESS_TOKEN)
xAI (Grok) grok XAI_API_KEY
Xiaomi (MiMo) Pay-as-you-go xiaomi_mimo_payg XIAOMI_MIMO_PAYG_API_KEY
Xiaomi (MiMo) Token Plan xiaomi_mimo_token_plan XIAOMI_MIMO_TOKEN_PLAN_API_KEY (+ region cn / sgp / ams)
Z.AI — Coding Plan zai_coding ZAI_CODING_API_KEY (+ region international / cn)
Z.AI — Pay-as-you-go zai_payg ZAI_API_KEY (+ region international / cn)
StepFun stepfun STEP_API_KEY (+ region global / cn)
OpenGateway opengateway OPENGATEWAY_API_KEY
Fireworks AI fireworks FIREWORKS_API_KEY

Runtime auto-detection uses a separate priority order (config.APIProviderDetectionOrder) when no deployment is pinned.

Usage

Basic Chat

resp, err := c.Chat(ctx, messages, llm.ChatOptions{
    Model: "gpt-4o",
})

Streaming with Continuation

// Auto-continues when max_tokens is hit
resp, err := provider.ChatWithContinuation(ctx, provider, messages,
    llm.ChatOptions{Model: model},
    core.DefaultContinuationConfig(),
)

Mock Provider for Testing

mock := provider.NewMockProvider(provider.MockModeFixed)
mock.Response = "Here is the code you asked for..."

resp, _ := mock.Chat(ctx, messages, opts)
// No real API calls — perfect for tests

Model Catalog

cat := catalog.DefaultModelCatalog()

// Get the best model for a tier
model := catalog.GetPreferredProviderModel("anthropic", catalog.TierSonnet, &cat)
// → "claude-sonnet-4-6"

// Check deprecation warnings
warn := catalog.GetModelDeprecationWarning("claude-3-7-sonnet", "anthropic")

Provider Configuration

cfg := config.LoadProviderConfig("")             // load from disk
config.ApplyProviderConfigToEnv(cfg, false, nil) // apply to environment
config.SaveProviderConfig(cfg, "")               // save changes

Architecture

flux/
├── engine/                 # Stable host-facing facade (hosts import engine, llm, graph, tools)
├── llm/                    # Host-facing DTOs and the Provider port that engine re-exports
├── graph/                  # Portable execution-graph vocabulary
├── tools/                  # Tool-call and tool-result contracts
├── provider/               # Provider runtime composition root (FluxClient)
│   ├── core/               # Provider-neutral wire, stream, retry, and transport primitives
│   ├── adapters/           # Provider protocol adapters and construction registry
│   ├── resilience/         # Rate limits, continuation, guardrails, and error policy
│   ├── cache/              # Response and semantic caches
│   ├── batch/              # Batch execution
│   ├── embeddings/         # Embedding clients, cache, and defaults
│   ├── media/              # Image and audio clients, structured prompts
│   ├── extraction/         # Structured extraction
│   ├── observability/      # Usage, cost, metrics, tracing, and recording
│   └── testkit/            # Mock provider for tests
├── catalog/                # Model catalog & tier system
│   ├── registry/           # Provider registry (single source of truth for providers)
│   ├── discover/           # Model discovery
│   ├── live/               # Live model listing per provider
│   ├── capabilities/       # Capability and deprecation data
│   └── concentrate/ opencodego/ opengateway/ xiaomi/ zai/  # Gateway-specific helpers
├── config/                 # Provider configuration & routing
│   └── credential/         # Credential file management
├── credentials/            # Keyring/env credential stores and OIDC keyless auth
├── router/                 # Routing strategies, deployment router, circuit breakers
│   └── controlplane/       # Versioned, signed peer manifests and replicas
├── runtime/                # Engine-internal provider/model/credential resolution
├── setup/                  # Catalog-backed deployment wiring
├── operationsgraph/        # Privacy-safe route and generation telemetry projection
├── conversation/           # Conversation engine with branching
├── storage/                # SQLite conversation DAG store, virtual keys, budgets
├── codeagent/              # Code agent retry & fallback strategies
├── verify/                 # Provider conformance harness
├── types/                  # Shared message types & API errors
├── constants/              # API limits
├── utils/                  # Error utilities
├── api/                    # OpenAPI spec for internal/api
├── internal/
│   ├── api/                # HTTP API server (library code; no flux binary starts it)
│   ├── cache/              # Cache backends and response cache warmer
│   ├── grpc/               # Optional gRPC transport (build tag grpc)
│   ├── health/             # Provider health checker
│   ├── httputil/ probehttp/ shrink/  # HTTP, probe, and tool-description helpers
│   ├── observability/      # OpenTelemetry spans, metrics, and audit sinks
│   └── sdk/                # Go, Python, TypeScript clients for the internal/api HTTP surface
├── docs/                   # Documentation & guides
├── examples/               # Runnable code examples
├── scripts/                # CI guards and helper scripts
└── assets/                 # Logo and branding

See docs/ARCHITECTURE.md for detailed system design and data flows.

operationsgraph.Build projects resolved routes and normalized usage into flux.graph/v1 operations nodes. Provider, model, request ID, and generated content are represented only by SHA-256 digests; token counts, finish reason, tool-call count, and deployment-routing state remain queryable.

Ecosystem

flux is part of the graycode-eco:

Component Repository Purpose
rho GrayCodeAI/rho AI coding agent
flux This repo LLM provider runtime

Development

Prerequisites

  • Go 1.26+

Build & Test

go build ./...               # Verify the library compiles
go test -race ./...           # Run all tests with race detector
make ci                       # Run full CI suite (lint, test, security)
make cover                    # Generate coverage report

Contributing

We welcome contributions! Please see CONTRIBUTING.md for development setup, commit conventions, and the PR process.

Quick start:

  1. Fork and create a branch: git checkout -b feat/short-description
  2. Make changes in small, focused commits
  3. Run make ci locally
  4. Open a pull request

Use Conventional Commits for commit messages — release-please uses them for versioning.

License

MIT — see LICENSE for details.

© 2026 GrayCode AI

About

A universal gateway for AI models.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages