diff --git a/docs/prepare-db-upgrade.md b/docs/prepare-db-upgrade-downgrade.md similarity index 51% rename from docs/prepare-db-upgrade.md rename to docs/prepare-db-upgrade-downgrade.md index 58806ecebebf..1ea1dcffd45d 100644 --- a/docs/prepare-db-upgrade.md +++ b/docs/prepare-db-upgrade-downgrade.md @@ -13,6 +13,10 @@ This document explains how to write or generate those scripts. 2. Run `misc/scripts/prepare-db-upgrade.sh --lang `. This will generate skeleton upgrade/downgrade scripts in the appropriate directories. 3. Fill in the details in the two `upgrade.properties` files that it generated, and add any required upgrade queries. +The generated directory names are hashes of the old and new `.dbscheme` files. If the +schema changes after generating the scripts, delete the generated directories and run the +script again so that the directory names and schema snapshots use the correct hashes. + It may be helpful to look at some of the existing upgrade/downgrade scripts, to see how they work. ## Details @@ -25,26 +29,52 @@ compatibility: partial some_relation.rel: run some_relation.qlo ``` -The `description` field is a textual description of the aim of the upgrade. +The `description` field is a textual description of the aim of the step. Describe the +operation in its actual direction: for example, a downgrade that removes a newly added +table should say that it removes the table. + +The `compatibility` field takes one of four values. In these definitions, the source schema is +`old.dbscheme`, and the target schema is the other `.dbscheme` in the script directory. Thus, +the source is the older schema for an upgrade and the newer schema for a downgrade. + + * **full**: query results from the transformed database will be identical to results from a database built with the target version of the toolchain. + + * **backwards**: the step is safe and preserves the meaning of the source database, but features provided by the target query and library packs may not work correctly on the transformed database. -The `compatibility` field takes one of four values: + * **partial**: the step is safe and preserves the meaning of the source database, but rebuilding the database with the target version of the toolchain would produce better results. - * **full**: results from the upgraded snapshot will be identical to results from a snapshot built with the new version of the toolchain. + * **breaking**: the step is unsafe and will prevent certain target queries from working. - * **backwards**: the step is safe and preserves the meaning of the old database, but new features may not work correctly on the upgraded snapshot. +Choose compatibility independently for the upgrade and downgrade, because the two directions +may preserve different amounts of information. - * **partial**: the step is safe and preserves the meaning of the old database, but you would get better results if you rebuilt the snapshot with the new version of the toolchain. +The `some_relation.rel` line(s) are the actions required to transform the database in the +direction of the step. Upgrade and downgrade directories use the same file name and command +syntax, even though a file in a downgrade directory describes a downgrade. Diff `old.dbscheme` +against the target `.dbscheme` in the generated directory to determine which actions are +needed. - * **breaking**: the step is unsafe and will prevent certain queries from working. +No action is needed for a relation added by the target schema if it should be empty in the +transformed database. A missing relation is treated as empty, so do not add a `.rel` line just +to create an empty file. If extraction would populate the new relation, however, the upgrade +is not `full`: it will usually be `backwards`, because queries using the new relation may have +degraded results on upgraded databases. -The `some_relation.rel` line(s) are the actions required to perform the database upgrade. Do a diff on the new vs old `.dbscheme` file to get an idea of what they have to achieve. Sometimes you won't need any upgrade commands – this happens when the dbscheme has changed in "cosmetic" ways, for example by adding/removing comments or changing union type relationships, but still retains the same on-disk format for all tables; the purpose of the upgrade script is then to document the fact that it's safe to replace the old dbscheme with the new one. +A relation that exists in `old.dbscheme` but not in the target schema should normally be +deleted explicitly with `relation.rel: delete` so that the transformation does not leave +obsolete data behind. + +Sometimes no commands are needed because the schema changed only cosmetically, for example +by adding or removing comments or changing union type relationships without changing the +on-disk format. The script then documents that it is safe to replace the old schema with the +new one. Ideally, your downgrade script will perfectly revert the changes applied by the upgrade script, such that applying the upgrade and then the downgrade will result in the same database you started with. -Some typical upgrade commands look like this: +Some typical upgrade or downgrade commands look like this: ``` -// Delete a relation that has been replaced in the new scheme +// Delete a relation that does not exist in the target schema obsolete.rel: delete // Create a new version of a table by applying an expression (using a simple @@ -83,7 +113,7 @@ To test the upgrade script, run: codeql test run --search-path= --search-path= ``` -Where `` is an extractor pack containing the old extractor and dbscheme that pre-date your changes, `` is the directory containing the qltests for your language, and `` is the root directory directory of the `github/codeql` clone that contains ``. This will run the tests using an old extractor, and the test databases will all be upgraded in place using your new upgrade script. +Where `` is an extractor pack containing the old extractor and dbscheme that pre-date your changes, `` is the directory containing the qltests for your language, and `` is the root directory of the `github/codeql` clone that contains ``. This will run the tests using an old extractor, and the test databases will all be upgraded in place using your new upgrade script. To test the downgrade script, create an extractor pack that includes your new dbscheme and extractor changes. Then checkout the `main` branch of `codeql` (i.e. a branch that does not include your changes), and run: @@ -107,43 +137,53 @@ You might also choose to test with a real-world database. 5. Verify that your queries produced sensible results. -#### Doing the upgrade manually +#### Creating the scripts manually -To create the upgrade directory manually, without using `prepare-db-upgrade.sh`: +To create both directions manually, without using `prepare-db-upgrade.sh`, run the following +commands from the repository root. First set `lang` to the language directory and `schema_file` +to the repository-relative path of its `.dbscheme` file. For example, for Go: -1. Get a hash of the old `.dbscheme` file from `main` (i.e. from just before your changes). You can do this by checking out the code prior to your changes and running `git hash-object ql/lib/.dbscheme` + ```sh + lang=go + schema_file=go/ql/lib/go.dbscheme + ``` -2. Go back to your branch and create an upgrade directory with that hash as its name, for example: -``` -mkdir ql/lib/upgrades/454f1e15151422355049dc4f1f0486a03baeffef -``` +1. Get the hashes of the old `.dbscheme` from `main` and the new `.dbscheme` from + your branch. For example: + ```sh + old_hash=$(git show "main:$schema_file" | git hash-object --stdin) + new_hash=$(git hash-object "$schema_file") + ``` -3. Copy the old `.dbscheme` file to that directory, using the name old.dbscheme. +2. Create the upgrade directory using the old hash and the downgrade directory using the + new hash: -``` -cp ql/lib/.dbscheme ql/lib/upgrades/454f1e15151422355049dc4f1f0486a03baeffef/old.dbscheme -``` - -4. Put a copy of your new `.dbscheme` file in that directory and create an `upgrade.properties` file (as described above). + ```sh + upgrade_dir="$lang/ql/lib/upgrades/$old_hash" + downgrade_dir="$lang/downgrades/$new_hash" + mkdir -p "$upgrade_dir" "$downgrade_dir" + ``` -#### Doing the downgrade manually +3. Populate the upgrade directory. Here, `old.dbscheme` is the schema from `main`, and + the other `.dbscheme` file is the new target schema: -The process is similar for downgrade scripts, but there is a reversal in terminology: your **new** dbscheme will now be the one called `old.dbscheme`. + ```sh + git show "main:$schema_file" > "$upgrade_dir/old.dbscheme" + cp "$schema_file" "$upgrade_dir/$(basename "$schema_file")" + ``` -1. Get a hash of your new `.dbscheme` file, with `git hash-object ql/lib/.dbscheme` +4. Populate the downgrade directory in the opposite direction. For a downgrade, the new + schema is called `old.dbscheme`, because it is the schema before the downgrade step: -2. Create a downgrade directory with that hash as its name, for example: -``` -mkdir downgrades/9fdd1d40fd3c3f8f9db8fabf5a353580d14c663a -``` - -3. Copy your new `.dbscheme` file to that directory, using the name `old.dbscheme`. -``` -cp ql/lib/.dbscheme ql/lib/upgrades/454f1e15151422355049dc4f1f0486a03baeffef/old.dbscheme -``` + ```sh + cp "$schema_file" "$downgrade_dir/old.dbscheme" + git show "main:$schema_file" > "$downgrade_dir/$(basename "$schema_file")" + ``` -4. Put a copy of the `.dbscheme` from `main` in that directory and create an `upgrade.properties` file that performs the downgrade (as described above). +5. Create an `upgrade.properties` file in each directory. The file in the upgrade directory + describes the forward transformation, while the file in the downgrade directory describes + the reverse transformation. ### Debugging your scripts