Remove terraform dependency in migrate tests - #6866
Merged
Merged
Conversation
The migrate tests deployed with the terraform engine to produce a terraform state, then migrated it to direct. Replace that live terraform deploy with a checked-in terraform.tfstate fixture: seed_tfstate.py deploys the bundle on the direct engine to create backend resources that match config (so the migrated state's first plan is a no-op), threads their ids into the fixture, and drops the direct state so the bundle resolves to the terraform engine again; the next deploy migrates it. The terraform engine is still present on this branch, so migration behavior is unchanged: a failed plan check still falls back to terraform (auto/plan-failure), and engine: terraform is still a deprecation warning (command/engine-config-terraform). Removing the engine, and adapting these two tests, is a later PR. Co-authored-by: Isaac <no-reply@databricks.com>
Collaborator
Integration test reportCommit: 9bcea5a
Top 6 slowest tests (at least 2 minutes):
|
Replace the direct-seed approach with record & replay for the permissions migrate test: check in the real terraform.tfstate + the terraform deploy's resource requests (captured from an actual terraform run), and replay the requests via `databricks api` so the backend holds real terraform-shaped resources, remapping recorded ids to the minted ones. The migrated state's out.original_state.json is now byte-identical to what terraform emits. Proof of the approach; the remaining migrate tests still use seed_tfstate.py and will be converted the same way. Co-authored-by: Isaac <no-reply@databricks.com>
Add capture_tfstate.py (freezes a real terraform deploy's tfstate + resource requests into the test source dir, normalizing host and UNIQUE_NAME to placeholders) and harden replay_tfstate.py to restore those placeholders for the current run. Convert auto/default to record/replay as the simple single-job example (permissions is the multi-resource one). Co-authored-by: Isaac <no-reply@databricks.com>
…riment resources Co-authored-by: Isaac <no-reply@databricks.com>
Committed 100644, so a fresh checkout (CI) fails every capture/replay with permission denied. Co-authored-by: Isaac <no-reply@databricks.com>
…istered_model/pipeline) Convert the 11 UC id-field/normalization migrate tests to the record/replay pattern: check in the real terraform.tfstate + the terraform deploy's resource requests (captured from a real terraform run), replay the requests so the backend holds terraform-shaped resources, and migrate. The recreate/rename/normalize behavior is preserved and now exercised against a real terraform state (e.g. reconcileIDFields sees the real deployed catalog_name). Also mark capture_tfstate.py and replay_tfstate.py executable (they were committed 100644, which fails with Permission denied on a fresh checkout / CI). Co-authored-by: Isaac <no-reply@databricks.com> (cherry picked from commit 0718bfbad797955212b096d5e3f92190e85d0a76)
Convert the auto fault/rejection migrate tests (dms-migration, plan-failure, plan-missing, push-failure, tfbackup-failure) off the terraform-engine deploy and onto replay_tfstate.py, matching the rest of the migrate suite. Add push_remote_tfstate.py: replay reconstructs the backend resources and the local terraform.tfstate, but a real terraform deploy also pushes the state to the workspace. Tests that inspect the remote state dir or back it up / delete it (tfbackup-failure's faulted delete, and the workspace-list assertions) push the replayed terraform.tfstate to the state path first, standing in for what a terraform deploy would have left. Co-authored-by: Isaac <no-reply@databricks.com> (cherry picked from commit 84ceabae402c35758eebff7fa57f0c1ca8b8c221)
Convert the simple auto-migrate tests from a live terraform deploy to the record/replay pattern. replay_tfstate.py replays the captured terraform requests against the backend and rewrites the checked-in terraform.tfstate ids to the freshly minted ones, so the migrate step reads a real, terraform-shaped state without the (soon-removed) terraform engine. Co-authored-by: Isaac <no-reply@databricks.com> (cherry picked from commit 65f518625be0d6321f958bd94cbb4f814a745528)
Replace the terraform-deploy step in the command/ migrate tests with
replay_tfstate.py (replaying a real terraform deploy's captured requests
against the fake backend, then dropping in the id-remapped tfstate), the
same approach already used for the auto/ tests.
Generalize the replay id-remapping so it also remaps backend-minted values
a later resource references - e.g. a volume's storage_location that a
pipeline tags via ${resources.volumes.*.storage_location}: the fake backend
mints a fresh storage_location on create, so without this the migrated
state keeps the recorded value and the first post-migration plan reports a
spurious pipeline update.
Drop the now-unused seed_tfstate.py.
Co-authored-by: Isaac <no-reply@databricks.com>
`databricks api --output json` decodes the response into Go's `any`, so JSON numbers come back as float64 and a freshly minted id above 2^53 (the test server mints ~8.7e18) loses its low digits. replay_tfstate.py read the minted id from that output, so it recorded a rounded id, wrote it into the tfstate, and the post-migration plan/deploy then looked the resource up by an id the backend never had - re-creating a "Migrated" resource instead of leaving it unchanged. It only bit under load (whether rounding changes the value depends on the id's low bits), so it surfaced as flaky migrate tests on the busy CI runners. Talk to the workspace directly with urllib instead: Python keeps the ids exact. Use a stable User-Agent so a test that records the replayed requests does not pin the Python version, and bypass the proxy (these tests only hit the local server). Also match a create to its tfstate resource by any identifying field, not just `name`: dashboards use `display_name`, so their id was never remapped and the follow-up publish call hit a stale id and failed - the dashboards test now performs a real migration instead of capturing that failure. Co-authored-by: Isaac <no-reply@databricks.com>
A real terraform deploy leaves its state both on disk and in the workspace. replay_tfstate.py wrote only the local copy, so a follow-up migrate had no remote state to back up and tests that inspect the remote state dir saw it empty - missing the terraform.tfstate.backup the migration produces (this had already been silently dropped from the artifacts and apply-failure tests). Push the remote copy from replay_tfstate.py itself rather than via a separate push_remote_tfstate.py that every test had to remember to call after replay. Removes push_remote_tfstate.py and its six call sites; the migrate tests that inspect the remote state dir now show the terraform.tfstate a real deploy would have left. Co-authored-by: Isaac <no-reply@databricks.com>
ruff format wraps the long workspace-import argument line; no behavior change. Co-authored-by: Isaac <no-reply@databricks.com>
…load The migrate tests keep their replay fixtures (terraform*.tfstate, terraform*-requests.json) and write out.* state dumps into the bundle root, so `bundle deploy` synced them as bundle files and inflated the uploaded file counts. script.prepare now writes a .gitignore - excluding itself too - so the deploy skips them, and test.toml adds .gitignore to Ignore so the harness does not flag it as an unexpected file. With the fixtures no longer uploaded there is no reason to delete them, so replay_tfstate.py is non-destructive; command/var_arg then replays one fixture twice instead of carrying a second identical copy. Co-authored-by: Isaac <no-reply@databricks.com>
denik
marked this pull request as ready for review
September 29, 2026 11:33
denik
enabled auto-merge
September 29, 2026 11:37
The Cloud=true migrate tests (schema/volume/catalog migration) replay against a real workspace, where auth is OAuth and DATABRICKS_TOKEN is empty - so replay's `Authorization: Bearer $DATABRICKS_TOKEN` came back HTTP 401 and every one of them failed on every cloud. When DATABRICKS_TOKEN is unset, resolve the token via `databricks auth token` (the same auth the bundle commands use); locally the PAT is still preferred so replay hits the same fake-server workspace as the bundle commands. Co-authored-by: Isaac <no-reply@databricks.com>
The migrate tests replay fixtures captured against the fake server, so they reproduce a terraform-shaped state only there. Eight of them were Cloud=true (on main they deployed on terraform against a real workspace), but replaying fake-server fixtures against a real workspace does not work: auth there is OAuth so the raw requests came back 401, and the real backend rejects mock-accepted bodies (catalog create 400). Mark them Cloud=false - they still exercise the migration logic against the fake server. Restoring cloud coverage for these scenarios is a follow-up, once replay can target a real workspace. Also revert the `databricks auth token` fallback in replay_tfstate.py: it was treating the symptom, and with replay local-only DATABRICKS_TOKEN is always the credential the fake server expects. Co-authored-by: Isaac <no-reply@databricks.com>
janniklasrose
approved these changes
Sep 29, 2026
github-merge-queue Bot
pushed a commit
that referenced
this pull request
Sep 30, 2026
`invariant/auto-migrate` and `invariant/migrate` seeded their migration by deploying on the terraform engine. So the engine can be removed, they now replay a per-config terraform-state fixture (captured once from a real terraform run) via `replay_tfstate.py`, like the `migrate/` tests. They run local-only, which is complete coverage rather than a gap: tf->direct migration is a one-time transition over a closed set of resource types (new resources are created on the direct engine and never migrated), so replaying the captured fixtures against the fake server covers it fully. Generalizes the capture/replay tooling for the fuzzer's resource diversity: - keep `workspace/mkdirs` (a dashboard's `parent_path`) and the recorded query params (postgres `?project_id=...`) in the captured fixtures; - match clusters by `cluster_name` / remap `cluster_id`; - for registered models and serving endpoints, skip the id-remap when the id is the identifying name they already sent, but still remap the backend-minted secondary id their permissions target (mlflow `registered_model_id`, serving endpoint id) -- for mlflow, fetched from the databricks registered-models GET the way terraform does. This lets `model_with_permissions` and `model_serving_endpoint` migrate. This pull request and its description were written by Isaac. Related: #6866 --------- Co-authored-by: Isaac <no-reply@databricks.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The migrate acceptance tests set up their starting state by running
DATABRICKS_BUNDLE_ENGINE=terraform bundle deploy. That stands in the way ofremoving the terraform engine.
Instead, each test now replays a real terraform deploy captured once as a
fixture: the resource-mutating requests (
terraform-requests.json) and theresulting
terraform.tfstate. At test timereplay_tfstate.pyreplays therequests against the fake backend, remaps the recorded ids (and other
backend-minted values, like a volume's
storage_location) to the freshlyminted ones, and drops in the id-remapped tfstate. Migrate then starts from a
real, terraform-shaped state without the terraform engine ever running.
Acceptance tests only; no product code changes.
This pull request and its description were written by Isaac.