Anonymized production case · C++

AsyncTaskManager

A subsystem that separated long-running operations from client requests and executed them in a pool of worker threads.

Environment
Production C++ service
Role
Design and implementation
Core
Queue, worker pool, task state
Synchronization
mutex, condition_variable
Task queue, task states and worker pool

Context

A client request should not wait for a long operation

The service had operations whose lifetime did not fit a synchronous request. Keeping the connection open would bind the client to processing time and make timeouts part of the operation contract.

Start separately

The client starts an operation and immediately receives a task identifier.

Check separately

Task state can be requested later by its identifier.

Execute in workers

Accepted operations enter a queue and are picked up by a fixed group of threads.

Keep state explicit

The task moves through queued, running and a final state instead of remaining hidden inside a request.

Execution

The lock protects the queue, not the work itself

A worker waits on a condition variable, removes one task while holding the mutex, updates its state and releases the lock. The operation then runs outside the critical section. Otherwise one long task would block other workers from taking work and the pool would effectively execute serially.

Submission and worker execution sequence with the operation running outside the lock
The protected section is limited to queue and state changes; the long operation runs after the mutex is released.

Shutdown

Stopping the service must also wake sleeping workers

The wait condition reacts both to new work and to a stop request. A worker exits only when shutdown has been requested and the queue is empty. This provides a conservative shutdown policy: stop accepting new work, finish tasks already accepted, then join the worker threads.

This shutdown policy is reconstructed from the component’s synchronization model. The case does not claim that a production deadlock or lost task occurred.

Failure handling

A failed operation should not kill a worker

The reconstructed design treats each operation as an isolated unit of work. An exception is caught at the worker boundary and converted into a failed task state, allowing the thread to continue processing the queue. The task identifier remains the client’s stable reference for both successful and failed completion.

Estimated operating range

The request returns quickly while work continues in the pool

The original measurements are no longer available. The ranges below are a conservative reconstruction for an in-memory queue and the described class of long-running operations, not benchmark results.

Task accepted

10–50 ms

Expected request time to validate the operation, enqueue it and return a task identifier.

Operation duration

5 s–several min

The approximate range in which separating work from the client request becomes useful.

Parallel execution

4–8 workers

A realistic pool size for CPU and I/O mixed work, with extra tasks waiting in the queue.

Client wait

100×+ shorter

For operations lasting several seconds, returning only the task identifier reduces request time by roughly two orders of magnitude.

Idle workers wait on the condition variable, so they do not continuously poll the queue and should add negligible idle CPU load.

Result

Request lifetime and operation lifetime became independent

The component gave the service a clear contract for starting long work, executing several operations concurrently and reporting their state.

Next project

Resume Assistant

A local application for preparing job applications without sending private data to the model.

Read case study →