Browse the catalogue: filter all 26 libraries by what they need, see real output, and get an install command for the headers you pick.
26 single-header C++17 libraries for building LLM features into native code. Streaming, retries, caching, cost estimation, RAG, reranking, tracing, structured output, agents and more. Each library is one .hpp file you copy into your project.
Most LLM tooling assumes Python or Node. If you are shipping a game, a desktop app, a trading system, an embedded tool or a C++ service, you usually end up hand-rolling HTTP calls, retry loops and JSON parsing. llm-cpp is that plumbing, split into small pieces so you take only what you need: no SDK, no package manager, no framework. The offline libraries have no dependencies at all; the ones that talk to OpenAI or Anthropic need only libcurl.
| I want to... | Use |
|---|---|
| Call a model and stream tokens | llm-stream |
| Build a chatbot with memory | llm-chat + llm-retry |
| Answer questions over my documents | llm-parse + llm-embed or llm-rag + llm-rank |
| Get valid JSON back every time | llm-format + llm-json |
| Let the model call my C++ functions | llm-agent |
| Know what my calls cost and where time goes | llm-cost + llm-log + llm-trace |
| Unit-test LLM code without the network | llm-mock |
"libcurl" means the implementation makes HTTPS calls (OpenAI and/or Anthropic APIs). "none" means it is fully offline and uses only the standard library.
| Library | What it does | Needs |
|---|---|---|
| llm-stream | Stream OpenAI and Anthropic chat responses token by token over SSE | libcurl |
| llm-retry | Exponential backoff with jitter, provider failover and a circuit breaker | none |
| llm-cost | Approximate token counts and cost estimates for built-in OpenAI and Anthropic models, budget checks | none |
| llm-cache | LRU response cache with TTL and hit/miss stats, so identical prompts skip the API | none |
| llm-format | Define a schema, validate model JSON against it, and re-prompt until the output conforms | none |
| llm-json | Small JSON parser and builder for request bodies and model output | none |
| Library | What it does | Needs |
|---|---|---|
| llm-parse | Strip HTML and markdown, extract titles, links, headings and code blocks, chunk text | none |
| llm-embed | OpenAI embeddings, cosine/dot/euclidean similarity and a small on-disk vector store | libcurl |
| llm-rag | End-to-end RAG: chunk, embed, persist an index, retrieve top-k and answer | libcurl |
| llm-rank | Rerank passages with offline BM25, LLM relevance scoring, or a hybrid of both | libcurl (linked; BM25 itself is offline) |
| llm-compress | Shrink conversation history: head/tail/smart truncation, sliding window, LLM summary | none (libcurl only with LLM_COMPRESS_SUMMARIZE) |
| llm-batch | Run a JSONL file of prompts through a thread pool with rate limiting and resumable checkpoints | libcurl |
| Library | What it does | Needs |
|---|---|---|
| llm-log | Structured JSONL log of every call with latency, tokens and cost, plus query and summary | none |
| llm-trace | RAII spans with parent/child nesting, token and cost attributes, OTLP-style JSON export | none |
| llm-pool | Worker pool with priority queue and requests-per-minute and tokens-per-minute limits | none |
| llm-mock | Fake LLM with scripted, pattern, random or echo responses, simulated latency and streaming | none |
| llm-eval | Run a prompt N times, measure consistency, compare models or prompts, score responses | libcurl |
| llm-ab | A/B test prompts or models with Welch's t-test, Cohen's d and custom scorers | libcurl |
| Library | What it does | Needs |
|---|---|---|
| llm-chat | Multi-turn conversation with token-budget trimming, pinned system prompt, save and restore | libcurl |
| llm-agent | Tool-calling agent loop: register C++ lambdas as tools and let the model call them | libcurl |
| llm-vision | Send images (file or URL) plus a prompt to OpenAI or Anthropic vision models | libcurl |
| llm-template | Mustache-style prompt templates with loops, conditionals and token-budget truncation | none |
| llm-router | Pick a model per prompt from a complexity score and a cost, latency, quality or budget strategy | none |
| llm-guard | Detect and scrub PII (email, phone, SSN, card numbers, API keys) and score prompt-injection risk | none |
| llm-audio | Whisper transcription and translation, and text-to-speech, via the OpenAI API | libcurl |
| llm-finetune | OpenAI fine-tuning lifecycle: write JSONL, upload, create, poll, cancel, list models | libcurl |
- Nothing to install.
curl -Oone file,#includeit. It works the same with CMake, Make, Bazel, MSBuild or a one-lineg++command. - You can read all of it. Each library is 210 to 572 lines; all 26 together are 8,923. When something misbehaves you open one file, not a dependency tree.
- You pay for what you use. Need retries and a cache? Take two headers. Nothing else is pulled in, and the offline ones add no link dependencies at all.
- Easy to vendor. Copy the headers into
third_party/, pin them in your own repo, patch them if you need to. No version resolver involved.
Grab the headers you want (each lives at include/<name>.hpp in its repo):
mkdir -p third_party && cd third_party
for lib in stream retry log; do
curl -fsSLO https://raw.githubusercontent.com/Mattbusel/llm-$lib/main/include/llm_$lib.hpp
doneEvery header follows the stb-style pattern: include it anywhere for the declarations, and in exactly one .cpp file define LLM_<NAME>_IMPLEMENTATION before including it to compile the implementation.
Six of the offline libraries have complete example programs in examples/offline, with the output they printed committed next to them. CI downloads each library's current header, builds every example with g++ and diffs the output, so these stay honest.
| Example | Shows |
|---|---|
| cache.cpp | LRU cache: case-insensitive hits, evictions, stats |
| cost.cpp | Price one prompt across the built-in models, block a call over budget |
| guard.cpp | Find and scrub PII and API keys, score prompt injection |
| format.cpp | Validate JSON against a schema and re-prompt until it conforms |
| json.cpp | Build a request body, read a response, reject bad input |
| compress.cpp | Keep a long chat inside a token budget with a sliding window |
curl -fsSLO https://raw.githubusercontent.com/Mattbusel/llm-guard/main/include/llm_guard.hpp
g++ -std=c++17 -I. examples/offline/guard.cpp -o guard && ./guardGive each implementation its own .cpp file. Several headers use the same internal helper names (for example llm::detail::json_escape), so defining two *_IMPLEMENTATION macros in one translation unit can fail to compile (llm-log with llm-stream is one such pair). In separate translation units they link together fine. As a check, all 26 implementations, each in its own .cpp, were compiled and linked into a single binary with MSVC 19.44 and libcurl on 2026-09-25.
// llm_impl_log.cpp
#define LLM_LOG_IMPLEMENTATION
#include "llm_log.hpp"
// llm_impl_retry.cpp
#define LLM_RETRY_IMPLEMENTATION
#include "llm_retry.hpp"
// llm_impl_stream.cpp
#define LLM_STREAM_IMPLEMENTATION
#include "llm_stream.hpp"Then use them together anywhere. This streams a completion, retries it on failure and writes a JSONL log line:
// main.cpp
#include "llm_log.hpp"
#include "llm_retry.hpp"
#include "llm_stream.hpp"
#include <cstdlib>
#include <iostream>
int main() {
const char* key = std::getenv("OPENAI_API_KEY");
if (!key) { std::cerr << "set OPENAI_API_KEY\n"; return 1; }
llm::Config cfg;
cfg.api_key = key;
cfg.model = "gpt-4o-mini";
const std::string prompt = "Explain backpressure in one paragraph.";
llm::Logger logger(llm::LogConfig{"calls.jsonl"});
llm::Logger::ScopedCall call(logger, cfg.model, prompt); // written on scope exit
auto result = llm::with_retry<std::string>([&]() -> std::string {
std::string text, error;
llm::stream(prompt, cfg,
[&](std::string_view tok) { std::cout << tok << std::flush; text += tok; },
nullptr,
[&](std::string_view err) { error = err; });
if (!error.empty()) throw llm::LLMError{0, error, true}; // retry
return text;
});
call.set_response(result.value);
std::cout << "\n(" << result.attempts_used << " attempt(s))\n";
}g++ -std=c++17 -O2 -Ithird_party main.cpp llm_impl_log.cpp llm_impl_retry.cpp llm_impl_stream.cpp -lcurl -o appAn offline pipeline needs no key and no network: clean a document with llm-parse, rank passages with llm-rank's BM25, and render the final prompt with llm-template.
#include "llm_parse.hpp"
#include "llm_rank.hpp"
#include "llm_template.hpp"
#include <iostream>
int main() {
std::string doc = llm::strip_html(
"<h1>Deploying</h1><p>Run make release to build the binary.</p>"
"<p>Copy config.yaml next to the binary.</p><p>Our office is in Berlin.</p>");
llm::ChunkConfig cc;
cc.chunk_size = 60;
cc.overlap = 0;
auto passages = llm::chunk(doc, cc);
std::string question = "how do I build the binary";
auto ranked = llm::rerank_local(question, passages);
llm::Template prompt("Answer using only this context:\n"
"{{#ctx}}- {{text}}\n{{/ctx}}\nQuestion: {{q}}\n");
llm::TemplateContext ctx;
ctx.vars["q"] = question;
for (size_t i = 0; i < ranked.size() && i < 2; ++i)
ctx.lists["ctx"].push_back({{"text", ranked[i].passage}});
std::cout << prompt.render(ctx);
}| Language | C++17 or later |
| Compilers | GCC, Clang, MSVC. Each library repo builds its examples in CI with CMake. |
| Network libraries | libcurl: preinstalled on macOS, apt install libcurl4-openssl-dev on Debian/Ubuntu, vcpkg install curl on Windows |
| Providers | OpenAI-compatible chat, embeddings, audio and fine-tuning endpoints; Anthropic Messages API in llm-stream and llm-vision |
These are small, focused libraries, not a full SDK. The HTTP code targets the public OpenAI and Anthropic endpoints and uses hand-written JSON handling, and token counts in llm-cost are approximations. Issues and pull requests are welcome in the individual repositories.
- LLMTokenStreamQuantEngine: C++20 engine that turns streaming LLM tokens into trade signals.
Need this kind of engineering on your product? I take on a small number of client builds: LLM features, iOS apps and performance work, fixed price. Services and pricing · Email · LinkedIn