Resync and reverify Geo data

  • Tier: Premium, Ultimate
  • Offering: GitLab Self-Managed

Resync and reverify are standard operational actions. You can use them to retry replication or verification without waiting for the automatic retry interval, or to recover after replication or verification failures.

In Rails console in a secondary Geo site, you can:

Resync and reverify individual components

On the secondary site, visit Admin > Geo > Replication to force a resync or reverify of individual items.

However, if this doesn’t work, you can perform the same action using the Rails console. The following sections describe how to use internal application commands in the Rails console to cause replication or verification for individual records synchronously or asynchronously.

Obtaining a Replicator instance

Commands that change data can cause damage if not run correctly or under the right conditions. Always run commands in a test environment first and have a backup instance ready to restore.

Before you can perform any sync or verify operations, you need to obtain a Replicator instance.

First, start a Rails console session in a primary or secondary site, depending on what you want to do.

Primary site:

  • You can checksum a resource

Secondary site:

  • You can sync a resource
  • You can checksum a resource and verify that checksum against the primary site’s checksum

Next, run one of the following snippets to get a Replicator instance.

Given a model record’s ID

ruby
model_record = Packages::PackageFile.find_by(id: 123)
replicator = model_record.replicator

Given a registry record’s ID

  • Replace 432 with the actual ID. A Registry record may or may not have the same ID value as the Model record that it tracks.
  • Replace Geo::PackageFileRegistry with any of the Geo Registry classes.

In a secondary Geo site:

ruby
registry_record = Geo::PackageFileRegistry.find_by(id: 432)
replicator = registry_record.replicator

Given an error message in a Registry record’s last_sync_failure

  • Replace Geo::PackageFileRegistry with any of the Geo Registry classes.
  • Replace error message here with the actual error message.
ruby
registry = Geo::PackageFileRegistry.find_by("last_sync_failure LIKE '%error message here%'")
replicator = registry.replicator

Given an error message in a Registry record’s verification_failure

  • Replace Geo::PackageFileRegistry with any of the Geo Registry classes.
  • Replace error message here with the actual error message.
ruby
registry = Geo::PackageFileRegistry.find_by("verification_failure LIKE '%error message here%'")
replicator = registry.replicator

Performing operations with a Replicator instance

After you have a Replicator instance stored in a replicator variable, you can perform many operations:

Sync in the console

This snippet only works in a secondary site.

This executes the sync code synchronously in the console, so you can observe how long it takes to sync a resource, or view a full error backtrace.

ruby
replicator.sync

Optionally, make the log level of the console more verbose than the configured log level, and then perform a sync:

ruby
Rails.logger.level = :debug

Checksum or verify in the console

This snippet works in any primary or secondary site.

In a primary site, it checksums the resource and stores the result in the main GitLab database. In a secondary site, it checksums the resource, compares it against the checksum in the main GitLab database (generated by the primary site), and stores the result in the Geo Tracking database.

This executes the checksum and verification code synchronously in the console, so you can observe how long it takes, or view a full error backtrace.

ruby
replicator.verify

Sync in a Sidekiq job

This snippet only works in a secondary site.

It enqueues a job for Sidekiq to perform a sync of the resource.

ruby
replicator.enqueue_sync

Verify in a Sidekiq job

This snippet works in any primary or secondary site.

It enqueues a job for Sidekiq to perform a checksum or verify of the resource.

ruby
replicator.verify_async

Get a model record

This snippet works in any primary or secondary site.

ruby
replicator.model_record

Get a registry record

This snippet only works in a secondary site because registry tables are stored in the Geo Tracking DB.

ruby
replicator.registry

Geo data type Model classes

A Geo data type is a specific class of data that is required by one or more GitLab features to store relevant data and is replicated by Geo to secondary sites.

  • Blob types:
    • Ci::JobArtifact
    • Ci::PipelineArtifact
    • Ci::SecureFile
    • LfsObject
    • MergeRequestDiff
    • Packages::PackageFile
    • PagesDeployment
    • Terraform::StateVersion
    • Upload
    • DependencyProxy::Manifest
    • DependencyProxy::Blob
  • Git Repository types:
    • DesignManagement::Repository
    • ProjectRepository
    • ProjectWikiRepository
    • SnippetRepository
    • GroupWikiRepository
  • Other types:
    • ContainerRepository

The main kinds of classes are Registry, Model, and Replicator. If you have an instance of one of these classes, you can get the others. The Registry and Model mostly manage PostgreSQL DB state. The Replicator knows how to replicate or verify the non-PostgreSQL data (file/Git repository/Container repository).

Geo Registry classes

In the context of GitLab Geo, a registry record refers to registry tables in the Geo tracking database. Each record tracks a single replicable in the main GitLab database, such as an LFS file, or a project Git repository. The Rails models that correspond to Geo registry tables that can be queried are:

  • Blob types:
    • Geo::CiSecureFileRegistry
    • Geo::DependencyProxyBlobRegistry
    • Geo::DependencyProxyManifestRegistry
    • Geo::JobArtifactRegistry
    • Geo::LfsObjectRegistry
    • Geo::MergeRequestDiffRegistry
    • Geo::PackageFileRegistry
    • Geo::PagesDeploymentRegistry
    • Geo::PipelineArtifactRegistry
    • Geo::ProjectWikiRepositoryRegistry
    • Geo::SnippetRepositoryRegistry
    • Geo::TerraformStateVersionRegistry
    • Geo::UploadRegistry
  • Git Repository types:
    • Geo::DesignManagementRepositoryRegistry
    • Geo::ProjectRepositoryRegistry
    • Geo::ProjectWikiRepositoryRegistry
    • Geo::SnippetRepositoryRegistry
    • Geo::GroupWikiRepositoryRegistry
  • Other types:
    • Geo::ContainerRepositoryRegistry

Resync and reverify multiple components

When component resources fail to sync or verify, you can trigger bulk actions to re-kick the replication queue. These actions reset the retry count and schedule time back to 0, causing the system to process the failed resources sooner rather than waiting up to 1 hour.

These actions don’t immediately process the resources. Instead, they re-queue the background jobs that handle synchronization and verification. The actual replication work happens asynchronously through the standard Geo replication process.

How resync and reverification works

When you trigger a resync or reverification action, the system marks matching records as pending. The Geo resync and reverification background workers pick up these records and process them according to normal queue priority. This mechanism allows you to expedite the processing of failed resources without immediately blocking on the operation.

It is not possible to reverify a record which is not successfully synced. Only a synced record can be verified.

It is possible to trigger bulk actions from the UI or from the Rails console.

From the UI

You can schedule a full resync of all resources of one component from the UI:

  1. In the upper-right corner, select Admin.
  2. In the left sidebar, select Geo > Sites.
  3. Under Replication details, select the desired component.

Resync resources for the selected component

  1. Select Resync all: this resets the status of all records for the selected resource, regardless of whether they are already synced or not.
  2. Select Resync all failed: this resets all records for which sync failed.

Reverify resources for the selected component

  1. Select Reverify all: this resets the status of all records for the selected resource, regardless of whether they are already verified or not.
  2. Select Reverify all failed: this resets all records for which verification failed, but sync is successful.

Reverify one component on all sites

If the primary site’s checksums are in question, then you need to make the primary site recalculate checksums. A “full re-verification” is then achieved, because after each checksum is recalculated on a primary site, events are generated which propagate to all secondary sites, causing them to recalculate their checksums and compare values. Any mismatch marks the registry as sync failed, which causes sync retries to be scheduled.

You can recalculate the primary site’s checksum from the UI:

  1. In the upper-right corner, select Admin.
  2. In the left sidebar, select Monitoring > Data management.
  3. Select the desired component in the dropdown list.
  4. Select Checksum all.

Resync all, Reverify all and Checksum all trigger an update of all resources, regardless of whether they are already synced or verified. It should not be executed when there are thousands of an object type in the instance (for example, CI Job Artifacts).

From the Rails console

Commands that change data can cause damage if not run correctly or under the right conditions. Always run commands in a test environment first and have a backup instance ready to restore.

The following sections describe how to use internal application commands in the Rails console to cause bulk replication or verification.

Sync all resources of one component that failed to sync

The following script:

  • Loops over all failed repositories.
  • Displays the Geo sync and verification metadata, including the reasons for the last failure.
  • Attempts to resync the repository.
  • Reports back if a failure occurs, and why.
  • Might take some time to complete. Each repository check must complete before reporting back the result. If your session times out, take measures to allow the process to continue running such as starting a screen session, or running it using Rails runner and nohup.

Run this script on the secondary Geo site.

ruby
Geo::ProjectRepositoryRegistry.failed.find_each do |registry|
   begin
     puts "ID: #{registry.id}, Project ID: #{registry.project_id}, Last Sync Failure: '#{registry.last_sync_failure}'"
     registry.replicator.sync
     puts "Sync initiated for registry ID: #{registry.id}"
   rescue => e
     puts "ID: #{registry.id}, Project ID: #{registry.project_id}, Failed: '#{e}'", e.backtrace.join("\n")
   end
end; nil

Reverify all resources that failed to checksum on the primary site

The system automatically reverifies all resources that failed to checksum on the primary site, but it uses a progressive backoff scheme to avoid an excessive volume of failures.

Optionally, for example if you’ve completed an attempted intervention, you can manually trigger reverification sooner:

  1. SSH into a GitLab Rails node in the primary site.

  2. Open the Rails console.

  3. Replacing LfsObject with any of the Geo data type Model classes, mark all resources as pending verification:

    ruby
    LfsObject.verification_state_table_class.where(verification_state: 3).each_batch do |relation|
      relation.update_all(verification_state: 0)
    end