Testing RTOS and Real-Time Constraints in Embedded Systems

Testing RTOS and Real-Time Constraints in Embedded Systems

Real-time operating systems add an entire category of bugs that don't exist in conventional software: timing bugs. A function that works correctly 99.9% of the time but misses its deadline 0.1% of the time will cause a system failure in a real-time application.

Testing RTOS-based systems requires verifying not just what the software does, but when.

Real-Time Requirements and Testing

A real-time system has deadlines — time constraints that, if violated, constitute a system failure. There are two categories:

Hard real-time: missing a deadline causes a catastrophic failure. A pacemaker that misses its timing constraint, an anti-lock braking system that responds too slowly, an industrial controller that doesn't actuate a valve in time. These systems fail when they're too slow, not just when they're incorrect.

Soft real-time: missing a deadline degrades quality but doesn't cause failure. Audio streaming that occasionally drops a frame, video decoding with occasional late frames, a UI that misses a refresh cycle. Failure is acceptable, but infrequent.

Your test suite must verify timing constraints, not just functional correctness. A function that returns the right value 100ms too late is a bug, not a success.

Measuring Task Timing

The foundation of RTOS testing is instrumenting task execution time. You need to know:

  • Worst-case execution time (WCET) for each real-time task
  • Task response time from trigger to completion
  • Interrupt service routine (ISR) latency
  • Time between task deadline and actual completion

Hardware Timer Instrumentation

The most accurate approach: use hardware timers to timestamp task entry and exit.

#include "FreeRTOS.h"
#include "task.h"

// DWT cycle counter (ARM Cortex-M)
#define DWT_CYCCNT  (*(volatile uint32_t *)0xE0001004)
#define DWT_CTRL    (*(volatile uint32_t *)0xE0001000)

void enable_cycle_counter(void) {
    DWT_CTRL |= 1;  // Enable CYCCNT
}

typedef struct {
    uint32_t start_cycles;
    uint32_t end_cycles;
    uint32_t max_cycles;
    uint32_t invocations;
} task_timing_t;

static task_timing_t sensor_task_timing = {0};

void sensor_task(void *pvParameters) {
    while (1) {
        ulTaskNotifyTake(pdTRUE, portMAX_DELAY);  // Wait for trigger
        
        sensor_task_timing.start_cycles = DWT_CYCCNT;
        
        // Actual task work
        read_sensor_data();
        process_sensor_data();
        transmit_results();
        
        sensor_task_timing.end_cycles = DWT_CYCCNT;
        
        uint32_t elapsed = sensor_task_timing.end_cycles - sensor_task_timing.start_cycles;
        if (elapsed > sensor_task_timing.max_cycles) {
            sensor_task_timing.max_cycles = elapsed;
        }
        sensor_task_timing.invocations++;
    }
}

uint32_t get_sensor_task_wcet_us(void) {
    return sensor_task_timing.max_cycles / (SystemCoreClock / 1000000);
}

Testing Deadline Compliance

#define SENSOR_TASK_DEADLINE_US  5000  // 5ms deadline

void test_sensor_task_meets_deadline(void) {
    // Reset timing data
    memset(&sensor_task_timing, 0, sizeof(sensor_task_timing));
    
    // Run the task 1000 times
    for (int i = 0; i < 1000; i++) {
        xTaskNotifyGive(sensor_task_handle);
        vTaskDelay(pdMS_TO_TICKS(10));  // Allow task to complete
    }
    
    uint32_t wcet_us = get_sensor_task_wcet_us();
    
    TEST_ASSERT_EQUAL_UINT32(1000, sensor_task_timing.invocations);
    TEST_ASSERT_LESS_THAN(SENSOR_TASK_DEADLINE_US, wcet_us);
    
    printf("Sensor task WCET: %lu µs (deadline: %d µs)\n", 
           wcet_us, SENSOR_TASK_DEADLINE_US);
}

Priority Inversion Testing

Priority inversion is one of the most insidious RTOS bugs: a low-priority task holds a resource needed by a high-priority task, blocking the high-priority task while a medium-priority task runs. The result is the high-priority task behaving as if it has the lowest priority.

The classic example: Mars Pathfinder's reset-inducing priority inversion from 1997.

Detecting Priority Inversion

// Test that priority inheritance prevents priority inversion
void test_priority_inheritance(void) {
    SemaphoreHandle_t mutex = xSemaphoreCreateMutex();
    
    uint32_t high_priority_completion_time = 0;
    uint32_t low_priority_completion_time = 0;
    
    // Low-priority task that holds the mutex
    TaskHandle_t low_task = create_task(
        "low_task", PRIORITY_LOW, mutex, &low_priority_completion_time
    );
    
    // Give low task time to acquire mutex
    vTaskDelay(pdMS_TO_TICKS(10));
    
    // High-priority task that needs the mutex
    TaskHandle_t high_task = create_task(
        "high_task", PRIORITY_HIGH, mutex, &high_priority_completion_time
    );
    
    // Wait for both to complete
    vTaskDelay(pdMS_TO_TICKS(500));
    
    // With priority inheritance: high-priority task completes first
    // Without priority inheritance: could be reversed (priority inversion)
    TEST_ASSERT_LESS_THAN(
        low_priority_completion_time, 
        high_priority_completion_time
    );
    
    vSemaphoreDelete(mutex);
}

Stress Testing for Priority Inversion

Run the priority inversion scenario repeatedly with varying timing:

void test_priority_inversion_stress(void) {
    const int NUM_ITERATIONS = 10000;
    int priority_violations = 0;
    
    for (int i = 0; i < NUM_ITERATIONS; i++) {
        // Run the scenario with random timing offsets
        int result = run_priority_inversion_scenario(rand() % 10);
        if (result == PRIORITY_VIOLATED) {
            priority_violations++;
        }
    }
    
    // Should never have priority inversion with proper mutex inheritance
    TEST_ASSERT_EQUAL_INT(0, priority_violations);
}

Interrupt Latency Testing

ISR latency — the time between an interrupt signal and the first instruction of the ISR executing — is critical for hard real-time systems. It must be bounded and tested.

volatile uint32_t interrupt_trigger_time;
volatile uint32_t isr_entry_time;

void EXTI0_IRQHandler(void) {
    isr_entry_time = DWT_CYCCNT;
    
    // ISR work
    handle_interrupt();
    
    EXTI->PR = EXTI_PR_PR0;  // Clear interrupt flag
}

void test_interrupt_latency(void) {
    uint32_t max_latency_cycles = 0;
    
    for (int i = 0; i < 1000; i++) {
        interrupt_trigger_time = DWT_CYCCNT;
        trigger_software_interrupt();  // EXTI0 via SWIER
        
        // Wait for ISR to complete
        while (isr_entry_time == 0);
        
        uint32_t latency = isr_entry_time - interrupt_trigger_time;
        if (latency > max_latency_cycles) {
            max_latency_cycles = latency;
        }
        isr_entry_time = 0;
    }
    
    uint32_t max_latency_ns = (max_latency_cycles * 1000) / (SystemCoreClock / 1000000);
    
    // STM32F4 @168MHz should have <300ns interrupt latency
    TEST_ASSERT_LESS_THAN(300, max_latency_ns);
    
    printf("Max interrupt latency: %lu ns over 1000 samples\n", max_latency_ns);
}

Race Condition Detection

RTOS race conditions occur when two tasks access shared memory without proper synchronization. They're non-deterministic and often only appear under specific timing conditions.

Systematic Race Condition Testing

Use timing manipulation to force race condition windows:

// Test for data race on sensor_value
volatile int32_t sensor_value = 0;

void writer_task(void *pvParameters) {
    for (int i = 0; i < 10000; i++) {
        // Non-atomic 32-bit write (two instructions on some architectures)
        sensor_value = 0xDEADBEEF;
        vTaskDelay(0);  // Yield to maximize race window
        sensor_value = 0;
    }
}

void reader_task(void *pvParameters) {
    int *tear_count = (int *)pvParameters;
    
    for (int i = 0; i < 10000; i++) {
        int32_t value = sensor_value;
        // A torn read will produce a value that's neither 0 nor 0xDEADBEEF
        if (value != 0 && value != (int32_t)0xDEADBEEF) {
            (*tear_count)++;
        }
        vTaskDelay(0);
    }
}

void test_sensor_value_requires_atomic_access(void) {
    int tear_count = 0;
    
    xTaskCreate(writer_task, "writer", 256, NULL, PRIORITY_HIGH, NULL);
    xTaskCreate(reader_task, "reader", 256, &tear_count, PRIORITY_HIGH, NULL);
    
    vTaskDelay(pdMS_TO_TICKS(1000));
    
    // If this fails, sensor_value needs atomic access (critical section or atomic type)
    TEST_ASSERT_EQUAL_INT_MESSAGE(0, tear_count,
        "Torn reads detected — sensor_value needs atomic or mutex-protected access");
}

Stack Overflow Testing

Stack overflows are a leading cause of RTOS failures. Test stack usage under worst-case conditions.

Stack High Water Mark Analysis

void test_all_task_stack_usage(void) {
    typedef struct {
        const char *name;
        TaskHandle_t handle;
        uint32_t stack_size;
        uint32_t min_free_words;  // Failure threshold
    } task_stack_info_t;
    
    task_stack_info_t tasks[] = {
        {"sensor_task",   sensor_task_handle,   512, 50},
        {"control_task",  control_task_handle,  256, 30},
        {"comms_task",    comms_task_handle,    1024, 100},
        {"monitor_task",  monitor_task_handle,  256, 30},
    };
    
    // Run system under maximum load for 60 seconds
    simulate_worst_case_load(60000);
    
    // Check high water marks for all tasks
    for (int i = 0; i < sizeof(tasks)/sizeof(tasks[0]); i++) {
        UBaseType_t high_water_mark = uxTaskGetStackHighWaterMark(tasks[i].handle);
        
        printf("Task '%s': %u words free (minimum: %u)\n",
               tasks[i].name, high_water_mark, tasks[i].min_free_words);
        
        TEST_ASSERT_GREATER_THAN_MESSAGE(
            tasks[i].min_free_words,
            high_water_mark,
            tasks[i].name
        );
    }
}

Intentional Stack Overflow Testing

Verify your stack overflow hook handles overflow correctly:

// Override FreeRTOS stack overflow hook
void vApplicationStackOverflowHook(TaskHandle_t xTask, char *pcTaskName) {
    overflow_detected = true;
    strncpy(overflow_task_name, pcTaskName, sizeof(overflow_task_name));
    // In production: log, reset, or enter safe state
}

void test_stack_overflow_detection(void) {
    overflow_detected = false;
    
    // Create a task that will overflow its stack
    xTaskCreate(stack_overflow_task, "overflow_test", 
                64,  // Deliberately too small
                NULL, PRIORITY_LOW, NULL);
    
    vTaskDelay(pdMS_TO_TICKS(500));
    
    TEST_ASSERT_TRUE_MESSAGE(overflow_detected, 
        "Stack overflow should have been detected");
    TEST_ASSERT_EQUAL_STRING("overflow_test", overflow_task_name);
}

Schedulability Analysis Testing

Verify that your task set is schedulable — that all tasks can meet their deadlines given the CPU budget.

For Rate-Monotonic Scheduling (RMS), the utilization bound for N tasks is: U = N(2^(1/N) - 1)

For N=3 tasks: U ≤ 78% For N=4 tasks: U ≤ 76%

typedef struct {
    const char *name;
    uint32_t period_ms;
    uint32_t wcet_us;
} task_spec_t;

void test_system_schedulability(void) {
    task_spec_t tasks[] = {
        {"control_loop",  10,  800},   // 10ms period, 0.8ms WCET = 8% utilization
        {"sensor_read",   50,  2000},  // 50ms period, 2ms WCET = 4% utilization
        {"comms_tx",      100, 5000},  // 100ms period, 5ms WCET = 5% utilization
        {"status_update", 1000, 10000} // 1s period, 10ms WCET = 1% utilization
    };
    int N = sizeof(tasks) / sizeof(tasks[0]);
    
    float total_utilization = 0;
    for (int i = 0; i < N; i++) {
        float ui = (float)tasks[i].wcet_us / (tasks[i].period_ms * 1000.0f);
        total_utilization += ui;
        printf("Task '%s': utilization = %.1f%%\n", tasks[i].name, ui * 100);
    }
    
    // RMS schedulability bound for N tasks
    float rms_bound = N * (pow(2.0f, 1.0f/N) - 1.0f);
    
    printf("Total utilization: %.1f%% (RMS bound: %.1f%%)\n",
           total_utilization * 100, rms_bound * 100);
    
    TEST_ASSERT_LESS_THAN_FLOAT(rms_bound, total_utilization);
}

Testing with QEMU and Hardware

RTOS timing tests need real hardware for accurate results — QEMU can't emulate cycle-accurate hardware timers for most platforms. But QEMU is useful for functional testing of RTOS behavior (task scheduling, semaphores, queues) where timing precision isn't required.

Structure your tests in two layers:

QEMU layer (CI-friendly): functional correctness — does the task produce the right output? Does the semaphore synchronize correctly? Does priority inheritance resolve correctly (functionally, not timing)?

Hardware layer (requires target hardware): timing correctness — does the task meet its deadline? What is the actual WCET? What is ISR latency?

Run QEMU tests in CI. Run hardware tests weekly or on release builds with hardware-in-the-loop (HIL) infrastructure.


HelpMeTest monitors real-time systems in production, tracking response times and availability with sub-minute interval checks that detect when your embedded systems miss their expected check-in windows. Start free.

Read more

Start now free