Hi Felipe,
Sorry for the long message! Tried to add enough detail so this post can be read in isolation. Please consider the implementation an experiment: it's only there as a way to discuss the debugger <> PAL contract, run tests, and try to poke holes or find rough edges.
Following up on the three lookup mechanisms we discussed, I now have an experimental Swift/LLDB implementation tested with real Embedded Swift tasks on a Raspberry Pi Pico 2’s two Cortex-M33 cores. Links at the end of the message.
This builds on the platform-deferred storage mechanism: the runtime delegates the storage description to the linked PAL, which selects how LLDB locates the current task.
Execution Contexts
All three mechanisms operate on an execution context, a platform entity whose current-task state we want to inspect: an OS or RTOS thread on a threaded system, or a hardware thread, such as a CPU core or hart, on a bare-metal system. An explicit (index, context_kind) pair is supplied by the platform debugger integration:
context_kind identifies the namespace: SOFTWARE_THREAD for an OS/RTOS thread, or HARDWARE_THREAD for a logical CPU or hart.
index is the platform-defined lookup key within that namespace. It identifies the execution context, not a Swift task.
For example, (1, HARDWARE_THREAD) could identify CPU 1, while (41, SOFTWARE_THREAD) could identify RTOS thread 41. For the helper-based mechanisms, the key could also be a thread-control-block pointer encoded as uintptr_t, provided the platform and debugger agree on its meaning and lifetime.
Lookup Mechanisms
There are three mechanisms, reflecting different storage guarantees a platform might provide:
| Mechanism |
Contract |
Scenario |
| Indexed storage |
Publish one process-wide table through a base address, count, stride, and expected context kind. LLDB computes the selected context’s slot address using the index and reads its current task pointer without executing target code. |
Useful on platforms with a small, fixed number of execution contexts, like CPU cores. |
| Slot-address helper |
A helper is called with (index, context_kind) to obtain a stable task-pointer slot address. That address is cached for the execution-context lifetime, and its contents are read on subsequent queries. |
Useful for abstracting access to TLS when threads are recognized but their storage layout isn't. |
| Current-task helper |
A helper is called with (index, context_kind) on every query, and LLDB uses the returned task pointer. No stable storage address is required. |
Useful as a catch-all for any other platform whose storage model doesn't fit the other optimized cases. |
Indexed storage
The first case is a small, known set of execution contexts, typically CPU cores on a bare-metal system. The platform maintains one process-wide table with a current-task pointer for each context and publishes its base address, count, and byte stride. LLDB computes the selected context’s slot address and reads it without executing target code.
Each entry contains an AsyncTask *, not a numeric task ID. The slot address remains fixed while its contents change as different tasks execute there. NULL means the context is idle; otherwise, LLDB obtains the task ID from the pointed-to task object.
The platform’s scheduler/runtime integration maintains these entries whenever the active Swift task changes, including nested job execution and restoration. A plain global array is sufficient.
Slot-address helper
The second case is a platform with threads and some native or emulated TLS mechanism that LLDB does not understand. An RTOS with its own threading implementation is a motivating example.
The platform provides a helper that takes (index, context_kind) and returns the address of that execution context’s current-task slot. LLDB calls it on a cache miss, caches the address, and reads the slot’s changing contents on subsequent queries.
This requires an explicit stability guarantee: the slot must remain at the same address and belong to the same execution context throughout its lifetime. Changing the task stored there or moving the thread between CPUs is fine. Relocating the slot is not. The cache tracks context lifetime and identity, including numeric ID reuse, but cannot reliably detect a slot that silently moves while its context remains alive.
A NULL helper return means unavailable and is not cached as valid. A valid slot containing NULL means idle.
I'm still working through how cache invalidation should work for this option. A reused thread ID should not retain an old cached address. I have some ideas that I'm weighing, but before getting deeper into that, I wanted to prioritize getting feedback on the three approaches.
Current-task helper
The third case is a platform that cannot expose a fixed table or guarantee a stable slot, but can answer which task belongs to an execution context.
Its helper takes the same pair and returns the task pointer on every query. This accommodates unusual storage arrangements, including platforms without conventional TLS, and prioritizes enabling task-aware debugging over the cost of repeated inferior calls. It also works when TLS exists but its storage is unsuitable for caching. NULL means idle; the prototype reserves UINTPTR_MAX for an unavailable lookup.
Calling Helpers
Both helpers execute in the target. Current LLDB schedules the call on the selected context’s backing LLDB thread, but this does not guarantee physical-CPU affinity for a software thread or preserve a whole-system snapshot.
The supplied pair is the authoritative lookup subject. Implementations must resolve it rather than substitute their current CPU or ambient TLS, unless the platform/debugger integration separately guarantees that they correspond. Calls require explicit opt-in, and helpers must not allocate, block, or acquire locks. The procedural fallback is flexible about storage, not about call safety.
Experimental Results
The hardware fixture uses CPicoSDK’s multicore scheduler and real Swift tasks. It moves a task from core 0 to core 1 and back, runs a second task, and checks repeated queries and idle contexts. All three mechanisms passed.
At a source breakpoint immediately after await Task.yield(), the same argument-free command produced these results:
# Unmodified debugger
(lldb) language swift task info
error: could not find the task address
# Patched debugger
(lldb) language swift task info
(UnsafeCurrentTask) current_task = id:1 flags:running {
address = 0x20015980
id = 1
enqueuePriority = .medium
parent = nil
children = {}
}
The stop at Tasks.swift:24 ran on core 1; after another yield, line 26 ran on core 0. Both identified the same task ID and address. With task-as-thread presentation enabled, both appeared as Task 1 with the same synthetic thread identity.
Indexed lookup executed no helpers and preserved registers and both heartbeat counters during inspection. The slot helper executed once per execution context, which here means once per core. The value helper executed on every query. Helper calls restored the calling integer registers, but the peer core advanced with this OpenOCD setup.
The LLDB test run passed 129 variants: 60 indexed-storage variants, 48 helper variants, and 21 existing GDB-remote tests. Software-context behavior is covered by mock remote-server tests, not an actual RTOS integration. ARM32 logical async unwinding and complete task enumeration remain outside this experiment’s scope.
The experimental implementation is on the Swift draft PR and LLDB draft PR. And there's also a gist with a use example, very high level.
Would love to get some thoughts. I hope this is heading in a direction that could eventually land. I think the value in this approach is generalizing the problem while also creating some optimization opportunities. Looking forward to feedback and ideas on how this could be evaluated further.