Feature Flag Infrastructure Testing: How to Validate Flag Systems at Scale
Most teams test their applications with feature flags. Few teams test their feature flag infrastructure itself. This is a gap — a broken flag system means features either always-on or always-off for all users, bypassing your entire deployment strategy.
This guide covers testing the flag infrastructure: the systems that store flags, evaluate rules, and deliver assignments to applications.
What Can Go Wrong with Flag Infrastructure
Feature flag systems have several failure modes that application-level tests don't catch:
Assignment inconsistency: A user gets variant A, clears cookies, visits again, and gets variant B. Your experiment is contaminated.
Evaluation latency: Flag evaluation adds 200ms to every request. Your "fast checkout" is now slow.
Cache staleness: You update a flag in the dashboard. It takes 5 minutes to propagate. Users see the old state during the propagation window.
SDK failure handling: Your flag service is temporarily unreachable. What does your SDK return? "Default to control" is correct. "Crash the application" is not.
Rule evaluation bugs: A targeting rule says "users in segment A" but evaluates as "users NOT in segment A" due to a logic inversion bug.
Impression logging gaps: 15% of flag evaluations aren't logged. Your A/B test data is missing a systematic chunk of exposures.
Each of these requires explicit testing. None of them show up in standard functional tests.
Testing Assignment Consistency
Assignments must be deterministic. The same user + flag combination must always return the same variant.
describe('Flag assignment consistency', () => {
it('returns the same variant for the same user ID across multiple calls', async () => {
const userId = 'user-consistency-test-123';
const flagName = 'test-consistency-flag';
const results = await Promise.all(
Array(10).fill(null).map(() => flags.evaluate(flagName, { userId }))
);
const uniqueVariants = new Set(results.map(r => r.variant));
expect(uniqueVariants.size).toBe(1);
});
it('returns the same variant after simulated cache miss', async () => {
const userId = 'user-cache-test-456';
const flagName = 'test-consistency-flag';
const firstResult = await flags.evaluate(flagName, { userId });
// Simulate cache miss
await flags.clearCache();
const secondResult = await flags.evaluate(flagName, { userId });
expect(firstResult.variant).toBe(secondResult.variant);
});
it('different users get different assignments (distribution test)', async () => {
const flagName = 'test-50-50-flag';
const userCount = 1000;
const results = await Promise.all(
Array(userCount).fill(null).map((_, i) =>
flags.evaluate(flagName, { userId: `user-${i}` })
)
);
const controlCount = results.filter(r => r.variant === 'control').length;
const treatmentCount = results.filter(r => r.variant === 'treatment').length;
// With 1000 users and 50/50 split, expect ~500 each ± 5%
expect(controlCount).toBeGreaterThan(450);
expect(controlCount).toBeLessThan(550);
expect(treatmentCount).toBeGreaterThan(450);
expect(treatmentCount).toBeLessThan(550);
});
});The distribution test is important: if your hashing algorithm is biased, you'll get 45/55 or 60/40 splits instead of 50/50. This directly affects A/B test validity.
Testing Evaluation Performance
Flag evaluation happens on every request. Performance testing is mandatory:
describe('Flag evaluation performance', () => {
it('single flag evaluation completes in under 5ms', async () => {
const start = performance.now();
await flags.evaluate('test-flag', { userId: 'perf-test-user' });
const duration = performance.now() - start;
expect(duration).toBeLessThan(5);
});
it('bulk evaluation of 10 flags completes in under 20ms', async () => {
const flagNames = Array(10).fill(null).map((_, i) => `test-flag-${i}`);
const start = performance.now();
await flags.evaluateAll(flagNames, { userId: 'perf-test-user' });
const duration = performance.now() - start;
expect(duration).toBeLessThan(20);
});
it('sustained load: 100 evaluations/second without degradation', async () => {
const evaluations = 100;
const durations: number[] = [];
for (let i = 0; i < evaluations; i++) {
const start = performance.now();
await flags.evaluate('test-flag', { userId: `user-${i}` });
durations.push(performance.now() - start);
// Maintain ~100 req/s pace
await new Promise(resolve => setTimeout(resolve, 10));
}
const p99 = durations.sort((a, b) => a - b)[Math.floor(evaluations * 0.99)];
expect(p99).toBeLessThan(50); // p99 under 50ms even under load
});
});If your flag evaluation p99 is 200ms, you've added 200ms to every user request. This is a problem worth catching before production.
Testing SDK Failure Handling
What happens when the flag service is unreachable?
describe('SDK resilience', () => {
it('returns default value when flag service is unreachable', async () => {
// Simulate service outage
server.close();
const result = await flags.evaluate('test-flag', {
userId: 'user-123',
defaultVariant: 'control'
});
expect(result.variant).toBe('control');
expect(result.reason).toBe('default_used');
});
it('does not throw when flag service returns 500', async () => {
server.use(
http.get('/api/flags/evaluate', () => {
return new HttpResponse(null, { status: 500 });
})
);
// Should not throw — should return default silently
await expect(
flags.evaluate('test-flag', { userId: 'user-123' })
).resolves.toMatchObject({ variant: 'control', reason: 'default_used' });
});
it('uses cached value when flag service is slow', async () => {
// First call populates cache
await flags.evaluate('test-flag', { userId: 'user-123' });
// Second call with service unavailable should use cache
server.close();
const result = await flags.evaluate('test-flag', { userId: 'user-123' });
expect(result.reason).toBe('cache');
});
it('SDK timeout is bounded', async () => {
server.use(
http.get('/api/flags/evaluate', async () => {
await delay(10000); // 10s delay
return new HttpResponse(JSON.stringify({ variant: 'treatment' }));
})
);
const start = Date.now();
await flags.evaluate('test-flag', { userId: 'user-123' });
const duration = Date.now() - start;
// SDK should timeout before 10s — should be under 1s
expect(duration).toBeLessThan(1000);
});
});The timeout test is critical. An SDK that waits indefinitely for a flag service response will hang your entire application when the flag service has an outage.
Testing Rule Evaluation Correctness
Flag targeting rules need their own test suite:
describe('Flag targeting rules', () => {
const flagConfig = {
name: 'premium-feature',
rules: [
{ condition: { plan: 'enterprise' }, variant: 'treatment', percentage: 100 },
{ condition: { plan: 'pro' }, variant: 'treatment', percentage: 50 },
],
defaultVariant: 'control'
};
it('enterprise users always get treatment', async () => {
const user = { userId: 'ent-user-1', attributes: { plan: 'enterprise' } };
const result = await flags.evaluate('premium-feature', user);
expect(result.variant).toBe('treatment');
});
it('free users always get control', async () => {
const user = { userId: 'free-user-1', attributes: { plan: 'free' } };
const result = await flags.evaluate('premium-feature', user);
expect(result.variant).toBe('control');
});
it('pro users get 50/50 split', async () => {
const proUsers = Array(100).fill(null).map((_, i) => ({
userId: `pro-user-${i}`,
attributes: { plan: 'pro' }
}));
const results = await Promise.all(
proUsers.map(user => flags.evaluate('premium-feature', user))
);
const treatmentCount = results.filter(r => r.variant === 'treatment').length;
expect(treatmentCount).toBeGreaterThan(35); // Expect ~50 ± 15
expect(treatmentCount).toBeLessThan(65);
});
it('rule evaluation is case-insensitive for plan names', async () => {
// Test defensive rule evaluation
const user = { userId: 'edge-1', attributes: { plan: 'ENTERPRISE' } };
const result = await flags.evaluate('premium-feature', user);
// Depends on your flag system behavior — document the expectation
expect(['treatment', 'control']).toContain(result.variant);
});
});Rule evaluation correctness tests are the most important in this list. A targeting rule that inverts its condition silently (e.g., plan === 'enterprise' evaluating as plan !== 'enterprise') is a production bug that's very hard to detect from application tests alone.
Testing Flag Propagation
How long does a flag change take to reach all instances of your application?
describe('Flag propagation timing', () => {
it('flag update propagates to all instances within 30 seconds', async () => {
const flagName = 'propagation-test-flag';
// Set flag to false
await flagAdmin.setFlag(flagName, { defaultVariant: 'control' });
// Record state across multiple instances
const instance1 = flags.createClient({ instanceId: 'instance-1' });
const instance2 = flags.createClient({ instanceId: 'instance-2' });
const instance3 = flags.createClient({ instanceId: 'instance-3' });
// Change the flag
await flagAdmin.setFlag(flagName, { defaultVariant: 'treatment' });
const maxWait = 30000; // 30 seconds
const pollInterval = 1000;
const start = Date.now();
while (Date.now() - start < maxWait) {
const [r1, r2, r3] = await Promise.all([
instance1.evaluate(flagName, { userId: 'test-user' }),
instance2.evaluate(flagName, { userId: 'test-user' }),
instance3.evaluate(flagName, { userId: 'test-user' }),
]);
if (r1.variant === 'treatment' && r2.variant === 'treatment' && r3.variant === 'treatment') {
const propagationTime = Date.now() - start;
console.log(`Flag propagated in ${propagationTime}ms`);
expect(propagationTime).toBeLessThan(maxWait);
return; // Test passes
}
await new Promise(resolve => setTimeout(resolve, pollInterval));
}
throw new Error(`Flag did not propagate within ${maxWait}ms`);
});
});This test verifies your flag system's SLA. If your flag system takes 5 minutes to propagate changes, teams need to know that — it affects how quickly they can respond to incidents by disabling a flag.
Testing Impression Logging
Impression logs are the foundation of A/B test analysis. Logging gaps corrupt your data.
describe('Impression logging', () => {
it('logs an impression on every flag evaluation', async () => {
const impressionsSpy = jest.spyOn(impressionLogger, 'log');
await flags.evaluate('test-flag', { userId: 'user-123' });
expect(impressionsSpy).toHaveBeenCalledOnce();
expect(impressionsSpy).toHaveBeenCalledWith({
flagName: 'test-flag',
userId: 'user-123',
variant: expect.any(String),
timestamp: expect.any(Number)
});
});
it('does not log duplicates for the same user+flag in a session', async () => {
const impressionsSpy = jest.spyOn(impressionLogger, 'log');
// Evaluate same flag twice in same session
await flags.evaluate('test-flag', { userId: 'user-123', sessionId: 'session-456' });
await flags.evaluate('test-flag', { userId: 'user-123', sessionId: 'session-456' });
// Should only log once per session
expect(impressionsSpy).toHaveBeenCalledOnce();
});
it('logging failure does not break flag evaluation', async () => {
jest.spyOn(impressionLogger, 'log').mockRejectedValue(new Error('Logger down'));
// Flag evaluation should succeed even if logging fails
await expect(
flags.evaluate('test-flag', { userId: 'user-123' })
).resolves.toBeDefined();
});
});The third test is the most important: your impression logger failing should never prevent the application from functioning. Logging is observability, not a critical path.
Integration Testing: Flags + Application
Beyond unit testing the flag infrastructure, test the integration between your flag system and application in a realistic environment:
test('full stack: flag controls which checkout flow user sees', async ({ page }) => {
// Set up: enable the new checkout flag for our test user
await flagAdmin.setUserVariant('checkout-flow', testUserId, 'treatment');
// Act: user visits checkout
await page.goto('/checkout');
// Assert: new checkout UI is rendered (not legacy)
await expect(page.locator('[data-testid="new-checkout-header"]')).toBeVisible();
await expect(page.locator('[data-testid="legacy-checkout-header"]')).not.toBeVisible();
// Cleanup
await flagAdmin.clearUserVariant('checkout-flow', testUserId);
});This end-to-end test validates the complete chain: flag service → SDK → application code → rendered UI. It catches integration failures that unit tests miss.
Observability Checklist for Flag Systems
Beyond testing, your flag system needs observability to be operational:
- Dashboard showing active flags and their current rollout percentages
- Real-time evaluation count by flag and variant
- Error rate for flag evaluations (should be near 0%)
- SDK version distribution (are all clients using current SDK?)
- Propagation lag monitoring (time between flag change and client receipt)
- Alert on flag evaluation error rate > 0.1%
- Alert on propagation lag > 60 seconds
Without observability, you find out your flag system is broken when an A/B test produces garbage data or a feature flag stops working in production. Observability lets you catch it first.
Summary
Feature flag infrastructure testing covers: assignment consistency (same user always gets same variant), evaluation performance (under 5ms per evaluation), SDK resilience (graceful degradation when service is unavailable), rule evaluation correctness (targeting logic does what the rules say), propagation timing (SLA for flag updates reaching all instances), and impression logging completeness (no gaps in experiment data). Most teams skip this entirely. The ones that don't ship experiments with reliable data and deployments with predictable rollback behavior.