Experiment rollouts and feature flags

Experiment rollout issue

Each experiment should have an experiment rollout issue to track the experiment from rollout through to cleanup and removal. The rollout issue is similar to a feature flag rollout issue, and is also used to track the status of an experiment.

Create the rollout issue from the Experiment Rollout issue template during implementation. Set its due date when you create it, to the planned end of the experiment run, and adjust it if the run slips. After the deadline, the issue must be resolved and either:

  • It was successful and the experiment becomes the new default.
  • It was not successful and all code related to the experiment is removed.

In either case, an outcome of the experiment should be posted to the issue with the reasoning for the decision. The scoped experiment:: label on the rollout issue tracks the experiment’s status, and the cleanup issue follows the outcome label. For the experiment:: labels, see the Growth experimentation handbook.

Experiment validation approach

Validate an experiment at different stages of the development lifecycle:

StageWhat to validateTools
Local development and CIEvent shape: which events fire for each variant, their category, action, and label, and the gitlab_experiment context.A feature spec that asserts a tracking journey contract. Optionally, Snowplow Micro.
StagingEach variant behaves as expected (user acceptance testing), and events are received.Experiment Dashboard, to force each variant and check event counts. Optionally, the Growth Experiment Event Validation Dashboard, which shows the same data.
ProductionEvents flow into Snowflake correctly.Experiment Dashboard. Optionally, the GLEX Experiment Analysis Dashboard, which shows the same data.

The feature spec proves the event shape before code review, and CI proves it again on every pipeline. Staging and production do not re-validate the shape. They confirm that events flow through the pipeline.

Before review, record the proof in the rollout issue. Include the contract and feature spec paths, the passing run, and the captured events.yml for each variant. The run writes these files. For more information, see Read the events a spec captured. The rollout issue template shows the expected format.

Staging user acceptance testing gates the production rollout. Event counts do not. The dashboard refreshes daily and events can take 24 hours or more to appear, so do not hold the rollout for them.

Turn off all experiments

When there is a case on GitLab.com that necessitates turning off all experiments, we have this control.

You can toggle experiments on GitLab.com using the gitlab_experiment feature flag.

This can be done via ChatOps:

  • disable: /chatops gitlab run feature set gitlab_experiment false
  • enable: /chatops gitlab run feature delete gitlab_experiment
  • This allows the default_enabled value of true in the YAML to be honored.

Notes on feature flags

We use the terms “enabled” and “disabled” here, even though it’s against our documentation style guide recommendations because these are the terms that the feature flag documentation uses.

You may already be familiar with the concept of feature flags in GitLab, but using feature flags in experiments is a bit different. While in general terms, a feature flag is viewed as being either on or off, this isn’t accurate for experiments.

Generally, off means that when we ask if a feature flag is enabled, it always returns false, and on means that it always returns true. An interim state, considered conditional, also exists. We take advantage of this trinary state of feature flags. To understand this conditional aspect: consider that either of these settings puts a feature flag into this state:

  • Setting a percentage_of_actors of any percent greater than 0%.
  • Enabling it for a single user or group.

Conditional means that it returns true in some situations, but not all situations.

When a feature flag is disabled (meaning the state is off), the experiment is considered inactive. You can visualize this in the decision tree diagram as reaching the first Running? node, and traversing the negative path.

When a feature flag is rolled out to a percentage_of_actors or similar (meaning the state is conditional) the experiment is considered to be running where sometimes the control is assigned, and sometimes the candidate is assigned. We don’t refer to this as being enabled, because that’s a confusing and overloaded term here. In the experiment terms, our experiment is running, and the feature flag is conditional.

When a feature flag is enabled (meaning the state is on), the candidate is always assigned.

We should try to be consistent with our terms, and so for experiments, we have an inactive experiment until we set the feature flag to conditional. After which, our experiment is then considered running. If you choose to “enable” your feature flag, you should consider the experiment to be resolved, because everyone is assigned the candidate unless they’ve opted out of experimentation.