Resync and reverify Geo data
- Tier: Premium, Ultimate
- Offering: GitLab Self-Managed
Resync and reverify are standard operational actions. You can use them to retry replication or verification without waiting for the automatic retry interval, or to recover after replication or verification failures.
In Rails console in a secondary Geo site, you can:
Resync and reverify individual components
On the secondary site, visit Admin > Geo > Replication to force a resync or reverify of individual items.
However, if this doesn’t work, you can perform the same action using the Rails console. The following sections describe how to use internal application commands in the Rails console to cause replication or verification for individual records synchronously or asynchronously.
Obtaining a Replicator instance
Commands that change data can cause damage if not run correctly or under the right conditions. Always run commands in a test environment first and have a backup instance ready to restore.
Before you can perform any sync or verify operations, you need to obtain a Replicator instance.
First, start a Rails console session in a primary or secondary site, depending on what you want to do.
Primary site:
- You can checksum a resource
Secondary site:
- You can sync a resource
- You can checksum a resource and verify that checksum against the primary site’s checksum
Next, run one of the following snippets to get a Replicator instance.
Given a model record’s ID
- Replace
123with the actual ID. - Replace
Packages::PackageFilewith any of the Geo data type Model classes.
model_record = Packages::PackageFile.find_by(id: 123)
replicator = model_record.replicatorGiven a registry record’s ID
- Replace
432with the actual ID. A Registry record may or may not have the same ID value as the Model record that it tracks. - Replace
Geo::PackageFileRegistrywith any of the Geo Registry classes.
In a secondary Geo site:
registry_record = Geo::PackageFileRegistry.find_by(id: 432)
replicator = registry_record.replicatorGiven an error message in a Registry record’s last_sync_failure
- Replace
Geo::PackageFileRegistrywith any of the Geo Registry classes. - Replace
error message herewith the actual error message.
registry = Geo::PackageFileRegistry.find_by("last_sync_failure LIKE '%error message here%'")
replicator = registry.replicatorGiven an error message in a Registry record’s verification_failure
- Replace
Geo::PackageFileRegistrywith any of the Geo Registry classes. - Replace
error message herewith the actual error message.
registry = Geo::PackageFileRegistry.find_by("verification_failure LIKE '%error message here%'")
replicator = registry.replicatorPerforming operations with a Replicator instance
After you have a Replicator instance stored in a replicator variable, you can perform many
operations:
Sync in the console
This snippet only works in a secondary site.
This executes the sync code synchronously in the console, so you can observe how long it takes to sync a resource, or view a full error backtrace.
replicator.syncOptionally, make the log level of the console more verbose than the configured log level, and then perform a sync:
Rails.logger.level = :debugChecksum or verify in the console
This snippet works in any primary or secondary site.
In a primary site, it checksums the resource and stores the result in the main GitLab database. In a secondary site, it checksums the resource, compares it against the checksum in the main GitLab database (generated by the primary site), and stores the result in the Geo Tracking database.
This executes the checksum and verification code synchronously in the console, so you can observe how long it takes, or view a full error backtrace.
replicator.verifySync in a Sidekiq job
This snippet only works in a secondary site.
It enqueues a job for Sidekiq to perform a sync of the resource.
replicator.enqueue_syncVerify in a Sidekiq job
This snippet works in any primary or secondary site.
It enqueues a job for Sidekiq to perform a checksum or verify of the resource.
replicator.verify_asyncGet a model record
This snippet works in any primary or secondary site.
replicator.model_recordGet a registry record
This snippet only works in a secondary site because registry tables are stored in the Geo Tracking DB.
replicator.registryGeo data type Model classes
A Geo data type is a specific class of data that is required by one or more GitLab features to store relevant data and is replicated by Geo to secondary sites.
- Blob types:
Ci::JobArtifactCi::PipelineArtifactCi::SecureFileLfsObjectMergeRequestDiffPackages::PackageFilePagesDeploymentTerraform::StateVersionUploadDependencyProxy::ManifestDependencyProxy::Blob
- Git Repository types:
DesignManagement::RepositoryProjectRepositoryProjectWikiRepositorySnippetRepositoryGroupWikiRepository
- Other types:
ContainerRepository
The main kinds of classes are Registry, Model, and Replicator. If you have an instance of one of these classes, you can get the others. The Registry and Model mostly manage PostgreSQL DB state. The Replicator knows how to replicate or verify the non-PostgreSQL data (file/Git repository/Container repository).
Geo Registry classes
In the context of GitLab Geo, a registry record refers to registry tables in the Geo tracking database. Each record tracks a single replicable in the main GitLab database, such as an LFS file, or a project Git repository. The Rails models that correspond to Geo registry tables that can be queried are:
- Blob types:
Geo::CiSecureFileRegistryGeo::DependencyProxyBlobRegistryGeo::DependencyProxyManifestRegistryGeo::JobArtifactRegistryGeo::LfsObjectRegistryGeo::MergeRequestDiffRegistryGeo::PackageFileRegistryGeo::PagesDeploymentRegistryGeo::PipelineArtifactRegistryGeo::ProjectWikiRepositoryRegistryGeo::SnippetRepositoryRegistryGeo::TerraformStateVersionRegistryGeo::UploadRegistry
- Git Repository types:
Geo::DesignManagementRepositoryRegistryGeo::ProjectRepositoryRegistryGeo::ProjectWikiRepositoryRegistryGeo::SnippetRepositoryRegistryGeo::GroupWikiRepositoryRegistry
- Other types:
Geo::ContainerRepositoryRegistry
Resync and reverify multiple components
When component resources fail to sync or verify, you can trigger bulk actions to re-kick the replication queue. These actions reset the retry count and schedule time back to 0, causing the system to process the failed resources sooner rather than waiting up to 1 hour.
These actions don’t immediately process the resources. Instead, they re-queue the background jobs that handle synchronization and verification. The actual replication work happens asynchronously through the standard Geo replication process.
How resync and reverification works
When you trigger a resync or reverification action, the system marks matching records as pending. The Geo resync and
reverification background workers pick up these records and process them according to normal queue priority.
This mechanism allows you to expedite the processing of failed resources without immediately blocking on the operation.
It is not possible to reverify a record which is not successfully synced. Only a synced record can be verified.
It is possible to trigger bulk actions from the UI or from the Rails console.
From the UI
You can schedule a full resync of all resources of one component from the UI:
- In the upper-right corner, select Admin.
- In the left sidebar, select Geo > Sites.
- Under Replication details, select the desired component.
Resync resources for the selected component
- Select Resync all: this resets the status of all records for the selected resource, regardless of whether they are already synced or not.
- Select Resync all failed: this resets all records for which sync failed.
Reverify resources for the selected component
- Select Reverify all: this resets the status of all records for the selected resource, regardless of whether they are already verified or not.
- Select Reverify all failed: this resets all records for which verification failed, but sync is successful.
Reverify one component on all sites
If the primary site’s checksums are in question, then you need to make the primary site recalculate checksums.
A “full re-verification” is then achieved, because after each checksum is recalculated on a primary site, events
are generated which propagate to all secondary sites, causing them to recalculate their checksums and compare values.
Any mismatch marks the registry as sync failed, which causes sync retries to be scheduled.
You can recalculate the primary site’s checksum from the UI:
- In the upper-right corner, select Admin.
- In the left sidebar, select Monitoring > Data management.
- Select the desired component in the dropdown list.
- Select Checksum all.
Resync all, Reverify all and Checksum all trigger an update of all resources, regardless of whether they are already synced or verified. It should not be executed when there are thousands of an object type in the instance (for example, CI Job Artifacts).
From the Rails console
Commands that change data can cause damage if not run correctly or under the right conditions. Always run commands in a test environment first and have a backup instance ready to restore.
The following sections describe how to use internal application commands in the Rails console to cause bulk replication or verification.
Sync all resources of one component that failed to sync
The following script:
- Loops over all failed repositories.
- Displays the Geo sync and verification metadata, including the reasons for the last failure.
- Attempts to resync the repository.
- Reports back if a failure occurs, and why.
- Might take some time to complete. Each repository check must complete
before reporting back the result. If your session times out, take measures
to allow the process to continue running such as starting a
screensession, or running it using Rails runner andnohup.
Run this script on the secondary Geo site.
Geo::ProjectRepositoryRegistry.failed.find_each do |registry|
begin
puts "ID: #{registry.id}, Project ID: #{registry.project_id}, Last Sync Failure: '#{registry.last_sync_failure}'"
registry.replicator.sync
puts "Sync initiated for registry ID: #{registry.id}"
rescue => e
puts "ID: #{registry.id}, Project ID: #{registry.project_id}, Failed: '#{e}'", e.backtrace.join("\n")
end
end; nilReverify all resources that failed to checksum on the primary site
The system automatically reverifies all resources that failed to checksum on the primary site, but it uses a progressive backoff scheme to avoid an excessive volume of failures.
Optionally, for example if you’ve completed an attempted intervention, you can manually trigger reverification sooner:
SSH into a GitLab Rails node in the primary site.
Open the Rails console.
Replacing
LfsObjectwith any of the Geo data type Model classes, mark all resources aspending verification:LfsObject.verification_state_table_class.where(verification_state: 3).each_batch do |relation| relation.update_all(verification_state: 0) end