Skip to content

Add FXMacroData macro provider - #16

Merged
luisleo526 merged 2 commits into
pineforge-4pass:mainfrom
fxmacrodata:fxmacrodata-macro-provider
Oct 1, 2026
Merged

luisleo526 merged 2 commits into
pineforge-4pass:mainfrom
fxmacrodata:fxmacrodata-macro-provider

Conversation

@roberttidball

Copy link
Copy Markdown
Contributor

Summary

Port of the FXMacroData integration from pineforge-4pass/pineforge-engine#82, as @luisleo526 suggested there.

  • add FxMacroDataProvider, a MacroDataProvider that turns FXMacroData announcements into MacroObservation records
  • one observation per distinct value a period has had, with released_at_ms from the record's announcement_datetime and vintage_at_ms from when that value became available
  • FXMacroData-specific fields (revisions, vintage_status, observed_at_ns, publication_time_status, pagination) stay inside the adapter; nothing is added to the normalized models
  • no new dependency: the default transport is urllib run in a thread, and an FxMacroDataTransport can be injected for offline tests or another HTTP client
  • add docs/providers/fxmacrodata.md, a "Macro providers" section in the catalog, mkdocs nav and API reference entries, and update the data-model line that said no macro provider ships

Disclosure: I maintain FXMacroData. It is a commercial API. Without a key only USD is available, limited to the most recent 90 days and delayed by 15 minutes; other currencies and full history need a key.

Release and vintage times

  • MacroObservation needs a release time and has no "unknown" value, so records whose announcement_datetime is null are skipped rather than given one.
  • Records whose release time FXMacroData derived from the series' usual lag instead of capturing it (release_time_assumed) are also skipped by default. include_assumed_release_times=True keeps them.
  • A captured snapshot's epoch repeats the original release time, so it is dated by observed_at_ns instead. Using epoch would put a later revision at the first print's timestamp, which is the lookahead the contract is meant to prevent. The trade-off is that history collected after the fact is only visible from its collection date.
  • Releases dated before their period (flash estimates) are skipped, since released_at_ms >= period_end_ms is enforced.

Skipped records are counted by reason in provider.last_skipped, so they are visible rather than silently dropped.

Safety and behavior

  • the key is sent only in the X-API-Key header, never in the URL, exception messages or repr
  • HTTP 429/5xx are retried max_retries times with backoff, honoring Retry-After; timeouts are explicit
  • pages are capped at 100 and followed through pagination.has_more / next_offset; a cursor that does not advance raises
  • the keyless tier's 90-day cut and withheld recent releases raise an FxMacroDataAccessWarning
  • the provider is not added to the registry: the registry and pineforge-backtest expect a MarketDataProvider

Validation

  • pytest — all passed except Docker-gated skips. tests/test_server.py was not run because I ran the suite without the dev extra: the sandbox I test in blocks httpx2 2.5.x over open advisories (GHSA-8xx6-hgc6-gc2m, GHSA-7mj9-2mp8-4m2p on httpcore2)
  • tests/test_fxmacrodata.py: 14 tests, offline, using a trimmed two-page fixture of real revisions=all responses
  • ruff check . and ruff format --check on the changed files
  • mypy src — no errors in the new module (the only errors were in server.py, from fastapi/pydantic not being installed)
  • mkdocs build --strict

@luisleo526 luisleo526 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for porting this over, and for the clear write-up on release versus vintage times. The snapshot-dated-by-observed_at_ns choice is the right one for the no-lookahead contract, and the offline fixture tests are exactly what we want. I ran the full suite, ruff, mypy, mkdocs strict and the build locally against main and everything the PR adds passes.

A few changes before merge:

  1. Tracking parameters. docs/providers/fxmacrodata.md line 5 links to the FXMacroData reference with utm_* query parameters. Please drop the query string so the docs link to the plain URL.
  2. Empty revisions list. A record with "revisions": [] currently produces no observation and no last_skipped entry (_vintages builds zero entries). Could you fall back to the record's own val in that case, the way revisions: null does, or at least count it under a skip reason?
  3. Redirects. UrllibTransport uses urlopen, whose default redirect handler forwards every header, including X-API-Key, to the redirect target and will follow a redirect to plain http://. Please build an opener whose redirect_request returns None, so a 3xx surfaces as an FxMacroDataHTTPError instead of being followed.
  4. Page count. The pagination loop ends only when the server stops saying has_more. Please add a ceiling on the number of pages (a constant or constructor argument) and raise FxMacroDataDataError when it is exceeded.
  5. Retry-After. The header value goes straight into asyncio.sleep; a server sending Retry-After: 86400 would park the caller for a day. Please clamp it (for example to 60 seconds) or fall back to the backoff schedule above a cap.

Smaller points, take or leave:

  • quote(currency) and quote(key) use the default safe="/", so a key containing / or .. becomes extra path segments. safe="" or a slug check would close that.
  • Server error text is passed into the exception message verbatim. Redacting the key from it would make the "never in exception messages" guarantee hold even if the upstream echoes it.
  • In docs/providers.md the new "Macro providers" section lands above the paragraph starting "CSV, SQLite, and SQLAlchemy share runtime schema discovery", which now reads as part of the macro section. Moving the section below that paragraph fixes it. The new page is also missing the breadcrumb line the other provider pages open with.
  • source is set to the publisher (BLS). Our other adapters put the adapter identity there (ccxt, sqlalchemy:<venue>#<table>). fxmacrodata or fxmacrodata:BLS would keep provenance consistent. Happy to hear if you see it differently.
  • UrllibTransport and the Retry-After and JSON-decoding helpers have no tests. A small test against a localhost http.server would cover headers, parsing and the redirect behaviour once changed. A test for the withheld_count warning would be good too.
  • The docs table of skipped records does not mention that publication_time_status: unverified records are kept. One line saying so, and that a planned time can precede the actual release, would help users judge the data.

We can't reach the live API from CI, so the review takes the field semantics (date as period end, epoch on later source vintages being their own publication time, val always appearing in revisions) on your word. If any of those does not hold in practice, the doc page is the place to say so.

@roberttidball

Copy link
Copy Markdown
Contributor Author

Thanks for the careful review. All of it is in 1b904c2.

  1. The docs link is now the plain URL.
  2. "revisions": [] now falls back to the record's own val, the same as null. Test: test_empty_revisions_fall_back_to_the_record_value.
  3. UrllibTransport uses an opener whose redirect_request returns None, so a 3xx raises FxMacroDataHTTPError and nothing is re-sent. Tested against a local http.server: one request to the original path, none to the target.
  4. New max_pages constructor argument (default 1000, so 100,000 rows) raises FxMacroDataDataError when exceeded.
  5. New max_retry_after_seconds (default 60). A larger Retry-After falls back to the backoff schedule.

Smaller points:

  • currency must be three letters and key must match [a-z0-9_]+, checked before any request. Both are also quoted with safe="".
  • The key is redacted from the server's error text.
  • The "Macro providers" section now sits below the tabular paragraph, and the page has the breadcrumb line.
  • source is now fxmacrodata:<publisher>, e.g. fxmacrodata:BLS.
  • Added localhost tests for the transport (headers, JSON, non-JSON body, redirect refusal, Retry-After) and one for the withheld_count warning.
  • The docs now say that unverified records are kept and that a planned time can come before the actual release.

On the field semantics, I checked the live keyless USD responses (16 series, 119 records, 226 revisions) against the spec:

  • date: holds for periodic series. It is the month end for monthly data and the week-ending date for weekly data. Two cases are not period ends: for policy_rate it is the decision date, and for daily series such as yields it is the observation date. The doc page now says so. No sampled release came before its date.
  • val in revisions: present and non-null on all 226.
  • epoch on later vintages: this did not fully hold, and I fixed the code.
    • Revisions tagged legacy_snapshot repeat the original release time in epoch, the same as captured_snapshot. They are now dated by observed_at_ns and skipped when that is absent.
    • On revisions with no vintage_status, epoch is sometimes the vintage's publication time and sometimes the time it was collected. In the sample it was never earlier than the release, and vintage_at_ms is clamped to released_at_ms anyway. The doc page explains this.
    • Test: test_legacy_snapshot_is_not_dated_by_its_epoch.
  • Quarterly series could not be checked without a key, because the keyless window filters by period date.

Checks: pytest (27 in tests/test_fxmacrodata.py, full suite green apart from the Docker-gated skips), ruff check ., ruff format --check src tests, mypy src (no errors outside server.py, where fastapi wasn't installed), mkdocs build --strict.

@luisleo526 luisleo526 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, Robert. Every point from the review is addressed in 1b904c2, and the follow-up went further than asked: the legacy_snapshot handling and the per-series date semantics came out of checking the live responses, which is exactly the kind of care the vintage contract needs. I re-ran the suite, ruff, mypy, mkdocs strict and the build against the new head, and probed the redirect refusal, page cap, Retry-After clamp and empty revisions handling independently; all hold. Merging. One small optional follow-up if you ever touch the tests again: clearing http_proxy in the local-server fixture would keep those three tests hermetic on machines with a proxy configured. Welcome aboard as the first macro provider in pineforge-data.

@luisleo526
luisleo526 merged commit f452d86 into pineforge-4pass:main Oct 1, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants