Skip to content

Remove terraform dependency in migrate tests - #6866

Merged
denik merged 15 commits into
mainfrom
denik/migration-tests-no-tf
Sep 29, 2026
Merged

denik merged 15 commits into
mainfrom
denik/migration-tests-no-tf

Conversation

@denik

@denik denik commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The migrate acceptance tests set up their starting state by running
DATABRICKS_BUNDLE_ENGINE=terraform bundle deploy. That stands in the way of
removing the terraform engine.

Instead, each test now replays a real terraform deploy captured once as a
fixture: the resource-mutating requests (terraform-requests.json) and the
resulting terraform.tfstate. At test time replay_tfstate.py replays the
requests against the fake backend, remaps the recorded ids (and other
backend-minted values, like a volume's storage_location) to the freshly
minted ones, and drops in the id-remapped tfstate. Migrate then starts from a
real, terraform-shaped state without the terraform engine ever running.

Acceptance tests only; no product code changes.

This pull request and its description were written by Isaac.

The migrate tests deployed with the terraform engine to produce a terraform state, then
migrated it to direct. Replace that live terraform deploy with a checked-in terraform.tfstate
fixture: seed_tfstate.py deploys the bundle on the direct engine to create backend resources
that match config (so the migrated state's first plan is a no-op), threads their ids into the
fixture, and drops the direct state so the bundle resolves to the terraform engine again; the
next deploy migrates it.

The terraform engine is still present on this branch, so migration behavior is unchanged: a
failed plan check still falls back to terraform (auto/plan-failure), and engine: terraform is
still a deprecation warning (command/engine-config-terraform). Removing the engine, and adapting
these two tests, is a later PR.

Co-authored-by: Isaac <no-reply@databricks.com>
@github-actions github-actions Bot added the DABs DABs related issues label Sep 28, 2026
@eng-dev-ecosystem-bot

eng-dev-ecosystem-bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Integration test report

Commit: 9bcea5a

Run: 36574757321

Env 🔄​flaky ✅​pass 🙈​skip Time
✅​ aws linux 276 53 5:47
✅​ aws windows 278 51 5:06
✅​ azure linux 275 53 5:39
🔄​ azure windows 1 276 51 3:44
✅​ gcp linux 276 53 5:31
✅​ gcp windows 278 51 3:28
Test Name azure windows
🔄​ TestSyncIncrementalFileOverwritesFolder 🔄​f
Top 6 slowest tests (at least 2 minutes):
duration env testname
5:04 aws windows TestAccept
4:07 aws linux TestAccept
4:03 azure linux TestAccept
3:52 gcp linux TestAccept
3:27 azure windows TestAccept
3:26 gcp windows TestAccept

denik and others added 8 commits September 28, 2026 18:05
Replace the direct-seed approach with record & replay for the permissions migrate test:
check in the real terraform.tfstate + the terraform deploy's resource requests (captured
from an actual terraform run), and replay the requests via `databricks api` so the backend
holds real terraform-shaped resources, remapping recorded ids to the minted ones. The
migrated state's out.original_state.json is now byte-identical to what terraform emits.

Proof of the approach; the remaining migrate tests still use seed_tfstate.py and will be
converted the same way.

Co-authored-by: Isaac <no-reply@databricks.com>
Add capture_tfstate.py (freezes a real terraform deploy's tfstate + resource requests into the
test source dir, normalizing host and UNIQUE_NAME to placeholders) and harden replay_tfstate.py
to restore those placeholders for the current run. Convert auto/default to record/replay as the
simple single-job example (permissions is the multi-resource one).

Co-authored-by: Isaac <no-reply@databricks.com>
…riment resources

Co-authored-by: Isaac <no-reply@databricks.com>
Committed 100644, so a fresh checkout (CI) fails every capture/replay with permission denied.

Co-authored-by: Isaac <no-reply@databricks.com>
…istered_model/pipeline)

Convert the 11 UC id-field/normalization migrate tests to the record/replay pattern: check in the
real terraform.tfstate + the terraform deploy's resource requests (captured from a real terraform
run), replay the requests so the backend holds terraform-shaped resources, and migrate. The
recreate/rename/normalize behavior is preserved and now exercised against a real terraform state
(e.g. reconcileIDFields sees the real deployed catalog_name).

Also mark capture_tfstate.py and replay_tfstate.py executable (they were committed 100644, which
fails with Permission denied on a fresh checkout / CI).

Co-authored-by: Isaac <no-reply@databricks.com>
(cherry picked from commit 0718bfbad797955212b096d5e3f92190e85d0a76)
Convert the auto fault/rejection migrate tests (dms-migration, plan-failure,
plan-missing, push-failure, tfbackup-failure) off the terraform-engine deploy
and onto replay_tfstate.py, matching the rest of the migrate suite.

Add push_remote_tfstate.py: replay reconstructs the backend resources and the
local terraform.tfstate, but a real terraform deploy also pushes the state to
the workspace. Tests that inspect the remote state dir or back it up / delete
it (tfbackup-failure's faulted delete, and the workspace-list assertions) push
the replayed terraform.tfstate to the state path first, standing in for what a
terraform deploy would have left.

Co-authored-by: Isaac <no-reply@databricks.com>
(cherry picked from commit 84ceabae402c35758eebff7fa57f0c1ca8b8c221)
Convert the simple auto-migrate tests from a live terraform deploy to the
record/replay pattern. replay_tfstate.py replays the captured terraform
requests against the backend and rewrites the checked-in terraform.tfstate
ids to the freshly minted ones, so the migrate step reads a real,
terraform-shaped state without the (soon-removed) terraform engine.

Co-authored-by: Isaac <no-reply@databricks.com>
(cherry picked from commit 65f518625be0d6321f958bd94cbb4f814a745528)
Replace the terraform-deploy step in the command/ migrate tests with
replay_tfstate.py (replaying a real terraform deploy's captured requests
against the fake backend, then dropping in the id-remapped tfstate), the
same approach already used for the auto/ tests.

Generalize the replay id-remapping so it also remaps backend-minted values
a later resource references - e.g. a volume's storage_location that a
pipeline tags via ${resources.volumes.*.storage_location}: the fake backend
mints a fresh storage_location on create, so without this the migrated
state keeps the recorded value and the first post-migration plan reports a
spurious pipeline update.

Drop the now-unused seed_tfstate.py.

Co-authored-by: Isaac <no-reply@databricks.com>
@denik denik changed the title Seed migrate tests from a terraform-state fixture Stop deploying terraform in migrate tests Sep 28, 2026
`databricks api --output json` decodes the response into Go's `any`, so JSON
numbers come back as float64 and a freshly minted id above 2^53 (the test
server mints ~8.7e18) loses its low digits. replay_tfstate.py read the minted
id from that output, so it recorded a rounded id, wrote it into the tfstate,
and the post-migration plan/deploy then looked the resource up by an id the
backend never had - re-creating a "Migrated" resource instead of leaving it
unchanged. It only bit under load (whether rounding changes the value depends
on the id's low bits), so it surfaced as flaky migrate tests on the busy CI
runners.

Talk to the workspace directly with urllib instead: Python keeps the ids
exact. Use a stable User-Agent so a test that records the replayed requests
does not pin the Python version, and bypass the proxy (these tests only hit
the local server).

Also match a create to its tfstate resource by any identifying field, not
just `name`: dashboards use `display_name`, so their id was never remapped and
the follow-up publish call hit a stale id and failed - the dashboards test now
performs a real migration instead of capturing that failure.

Co-authored-by: Isaac <no-reply@databricks.com>
@denik denik changed the title Stop deploying terraform in migrate tests Remove terraform dependency in migrate tests Sep 29, 2026
denik and others added 3 commits September 29, 2026 12:53
A real terraform deploy leaves its state both on disk and in the workspace.
replay_tfstate.py wrote only the local copy, so a follow-up migrate had no
remote state to back up and tests that inspect the remote state dir saw it
empty - missing the terraform.tfstate.backup the migration produces (this had
already been silently dropped from the artifacts and apply-failure tests).

Push the remote copy from replay_tfstate.py itself rather than via a separate
push_remote_tfstate.py that every test had to remember to call after replay.
Removes push_remote_tfstate.py and its six call sites; the migrate tests that
inspect the remote state dir now show the terraform.tfstate a real deploy
would have left.

Co-authored-by: Isaac <no-reply@databricks.com>
ruff format wraps the long workspace-import argument line; no behavior change.

Co-authored-by: Isaac <no-reply@databricks.com>
…load

The migrate tests keep their replay fixtures (terraform*.tfstate,
terraform*-requests.json) and write out.* state dumps into the bundle root, so
`bundle deploy` synced them as bundle files and inflated the uploaded file
counts. script.prepare now writes a .gitignore - excluding itself too - so the
deploy skips them, and test.toml adds .gitignore to Ignore so the harness does
not flag it as an unexpected file.

With the fixtures no longer uploaded there is no reason to delete them, so
replay_tfstate.py is non-destructive; command/var_arg then replays one fixture
twice instead of carrying a second identical copy.

Co-authored-by: Isaac <no-reply@databricks.com>
@denik
denik marked this pull request as ready for review September 29, 2026 11:33
@denik
denik requested review from a team as code owners September 29, 2026 11:33
@denik
denik enabled auto-merge September 29, 2026 11:37
The Cloud=true migrate tests (schema/volume/catalog migration) replay against a
real workspace, where auth is OAuth and DATABRICKS_TOKEN is empty - so replay's
`Authorization: Bearer $DATABRICKS_TOKEN` came back HTTP 401 and every one of
them failed on every cloud. When DATABRICKS_TOKEN is unset, resolve the token
via `databricks auth token` (the same auth the bundle commands use); locally
the PAT is still preferred so replay hits the same fake-server workspace as the
bundle commands.

Co-authored-by: Isaac <no-reply@databricks.com>
Comment thread acceptance/bin/capture_tfstate.py
Comment thread acceptance/bundle/migrate/auto/apply-failure/terraform.tfstate
Comment thread acceptance/bundle/migrate/auto/artifacts/output.txt
The migrate tests replay fixtures captured against the fake server, so they
reproduce a terraform-shaped state only there. Eight of them were Cloud=true
(on main they deployed on terraform against a real workspace), but replaying
fake-server fixtures against a real workspace does not work: auth there is OAuth
so the raw requests came back 401, and the real backend rejects mock-accepted
bodies (catalog create 400). Mark them Cloud=false - they still exercise the
migration logic against the fake server. Restoring cloud coverage for these
scenarios is a follow-up, once replay can target a real workspace.

Also revert the `databricks auth token` fallback in replay_tfstate.py: it was
treating the symptom, and with replay local-only DATABRICKS_TOKEN is always the
credential the fake server expects.

Co-authored-by: Isaac <no-reply@databricks.com>
@denik
denik added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit e41a5c8 Sep 29, 2026
30 checks passed
@denik
denik deleted the denik/migration-tests-no-tf branch September 29, 2026 14:04
github-merge-queue Bot pushed a commit that referenced this pull request Sep 30, 2026
`invariant/auto-migrate` and `invariant/migrate` seeded their migration
by deploying on the terraform engine. So the engine can be removed, they
now replay a per-config terraform-state fixture (captured once from a
real terraform run) via `replay_tfstate.py`, like the `migrate/` tests.
They run local-only, which is complete coverage rather than a gap:
tf->direct migration is a one-time transition over a closed set of
resource types (new resources are created on the direct engine and never
migrated), so replaying the captured fixtures against the fake server
covers it fully.

Generalizes the capture/replay tooling for the fuzzer's resource
diversity:
- keep `workspace/mkdirs` (a dashboard's `parent_path`) and the recorded
query params (postgres `?project_id=...`) in the captured fixtures;
- match clusters by `cluster_name` / remap `cluster_id`;
- for registered models and serving endpoints, skip the id-remap when
the id is the identifying name they already sent, but still remap the
backend-minted secondary id their permissions target (mlflow
`registered_model_id`, serving endpoint id) -- for mlflow, fetched from
the databricks registered-models GET the way terraform does. This lets
`model_with_permissions` and `model_serving_endpoint` migrate.

This pull request and its description were written by Isaac.

Related: #6866

---------

Co-authored-by: Isaac <no-reply@databricks.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

DABs DABs related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants