Concurrency-aware debugging with Embedded Swift platform runtimes

Hello all,

Following the earlier discussion with @felipepiovezan in the Multicore Concurrency thread and his follow-up on swiftlang/swift#91047, I wanted to split Embedded Swift debugger support into a dedicated discussion.

The original question was how LLDB can determine which Swift task is executing on each thread. Since then, the Embedded Concurrency implementation has changed in a way that makes the question more general.

Embedded now builds Concurrency against the Embedded Threading backend. The runtime’s reserved TLS entries, including the current ActiveTask, go through Swift’s ThreadLocalStorage abstraction and ultimately through the swift_tls* platform hooks. The concrete storage behind those hooks is selected by the EmbeddedPlatform implementation linked into the final program.

That implementation might use:

  • native pthread thread-specific storage
  • dynamically allocated pthread keys
  • RTOS-provided task-local storage
  • a fixed per-core table
  • a single global slot for a single-threaded system
    or a completely unknown platform-specific alternative.

This creates a mismatch with the current debug ABI. _swift_concurrency_debug_internal_layout_version includes a _concurrency_current_task_storage_kind, which tells LLDB how the runtime stores the current task. The enum already recognizes that there are several possible storage strategies, but the Embedded Concurrency library cannot fully determine that strategy when it is built. It only knows that it will call the platform abstraction; the concrete representation is selected later at link time.

I think it is useful to separate this into two problems:

What is an execution context?
On a hosted platform this is usually an OS thread already represented by LLDB. On embedded targets it might instead be an RTOS task, a CPU core, or the only execution context in a single-threaded program. Some platforms may therefore require debugger or debug-server support before Swift can even answer questions about tasks.

Given an execution context, where is its current Swift task pointer?
Once LLDB has an appropriate thread or execution-context representation, it needs a fast and passive way to locate the corresponding ActiveTask storage.

In our earlier discussion, I suggested letting LLDB call a runtime or platform accessor. Felipe pointed out an important constraint:

This mapping of Threads -> Task has to be fast. It happens on every single stop [...] This is particularly important when communication between the debugger and the target is slower than local.

That makes an inferior call through _swift_tls_get unattractive as the general solution. It would require executing target code in the context of each thread, could perturb target state, and may not be available on bare-metal or remote targets.

Pausing here, hope this provides enough context to start a conversation.

Cheers.

3 Likes

Thank you for the summary, Gonzalo!

It's a sad common theme that most swift features get merged without a debugging story, but here we are; let's figure out what's the best way of remedying this.

I would like to offer one further way of looking at the current challenge; after the related changes got merged, the concurrency library can sometimes describe the thread local storage implementation used: it can do so for non-embedded cases, but it cannot do so for embedded cases.

I think it is useful to separate this into two problems:

What is an execution context?

While this is a real problem, I don't believe it is related to the problem at hand. The mapping of system-level abstractions into the "thread" abstraction LLDB presents to users is a well known problem, and lives on a different layer of the debugger. I'd like to make the assumption LLDB can do this for whatever system is being debugged (if it can't, then it can't debug any process for that system), and focus on the problems that surface once the mapping is complete.

Given an execution context, where is its current Swift task pointer?

As you summarised, LLDB made the assumption that the concurrency library fully describes everything about concurrency, which is no longer true. I see two ways around this:

  • Change the non-embedded case, so that we have a uniform representation.
  • Keep the non-embedded case as is, and add a new level of indirection for embedded: set the _concurrency_current_task_storage_kind to a new value called definedInShim, and create a new debug contract for those shims: they must expose some global symbol defining their TLS implementation. LLDB would then know to query a different symbol.

I'm leaning towards the second approach.

To answer something you proposed on the other thread:

Would it make sense to start with a PR to add an unavailable storage kind?

I think this would make the problem worse. The shim libswiftEmbeddedPlatformMultiThreadedDarwin uses the exact same implementation as the non embedded case on Darwin, so we patched LLDB to just assume that shim until we have a better story. Using unavailable would regress that case.

1 Like

Also, please let me know if I misunderstood what you meant by execution context!

Thanks, Felipe! Hopefully I can help solve this as much as I helped exacerbate the problem :sweat_smile:

Also, please let me know if I misunderstood what you meant by execution context!

Let me start with this. I’d define a platform execution context as the independently executing platform state associated with context-local storage: an OS thread on hosted systems, a hardware core or hart on bare-metal systems, or a scheduler-managed context on an RTOS.

I would like to offer one further way of looking at the current challenge; after the related changes got merged, the concurrency library can sometimes describe the thread local storage implementation used: it can do so for non-embedded cases, but it cannot do so for embedded cases.

Yes, thanks for the clarification. I was scoping my wording to Embedded debuggability, but this is a more accurate characterization of the current situation.

What is an execution context?

While this is a real problem, I don't believe it is related to the problem at hand. The mapping of system-level abstractions into the "thread" abstraction LLDB presents to users is a well known problem, and lives on a different layer of the debugger.

I have a slightly different view here, although I think the difference may mostly be about layering.

The way I understand it, obtaining the current task pointer is a function of the debugger-visible execution context. The existing implementations work because LLDB knows how to map an LLDB Thread to either compiler TLS or Darwin TSD. On a bare-metal target, a debugger might expose each core as an LLDB Thread, while the PAL stores current-task pointers in a CPU-indexed table. Some layer still needs to define how that debugger thread maps to the corresponding CPU index.

I agree that this mapping does not necessarily belong in the Swift Concurrency plugin. We can assume that the platform integration has already exposed the relevant contexts as LLDB threads and focus this contract on locating the current task once that mapping exists. I just want to make the prerequisite explicit, because without it the Embedded debugging story might remain incomplete.

Keep the non-embedded case as is, and add a new level of indirection for embedded: set the _concurrency_current_task_storage_kind to a new value called definedInShim, and create a new debug contract for those shims: they must expose some global symbol defining their TLS implementation. LLDB would then know to query a different symbol.

I was exploring this same approach, and I think it fits well. Concurrency explicitly declares that the decision has been delegated, and only then does LLDB consult the shim-provided value. That avoids the ambiguity of looking opportunistically for an optional secondary symbol.

Along with that, I’ve been exploring a few lookup options that a shim could select:

  • A fixed indexed-storage fast path for platforms with a small, stable set of execution contexts, such as CPU cores. LLDB could read the task pointer directly from a known table once it has the context index.
  • A helper that returns the address of the current context’s task-pointer slot. LLDB would call it once per debugger-visible thread and cache the address, assuming that address remains stable for the lifetime of the thread.
  • A fully general helper that returns the current task pointer and must be called on each lookup. This would support platforms that cannot expose stable storage, with the inferior-call performance cost you described earlier.

I think these options could give platforms a useful range of contracts, from direct memory access to a fully procedural lookup, depending on what their execution model can support. Does that sound like a reasonable shape for the shim-side debug contract?

I think I agree with everything you said.

Would you (or myself, either way is fine) like to give it a try to expose the definedInShim enum value and have the "multiple threads" shim define its "static tls key" entry defined? It would be effectively an NFC patch that lays out the foundation for other shims.

2 Likes

Awesome!

Sure, I can take care of that. Let me send you a PR later today / tomorrow.

1 Like

Alright, sorry this took a bit longer.

I put up two options. I’m leaning towards the first, but it’s a broader change, and I don’t want to stretch the original storage-kind field’s meaning too much :)

Option 1: Deferring to the platform library becomes a flag, separate from the concrete storage kinds. LLDB follows that indirection once and rejects a platform value that tries to defer again. The storage-kind definitions also move into a shared C-compatible header so platform implementers can use named ABI values.

Option 2: A more conservative approach that adds a platform_defined storage kind to indicate the indirection. It keeps the definitions in the existing debug header, with the PAL publishing the corresponding numeric constant.

Here are the LLDB changes for Option 1. If Option 2 is a better fit, adapting the decoder should be a small change.

Open to thoughts and feedback!

Thank you for putting those together, @Gonzalo_Larralde ! I like option 1 as well, and left some minor points in the two PRs.

Next step is to change the destination branch for the LLDB code to stable/21.x and run testing on those. I can do some local testing as well

(sorry for the delay, I was on vacation!)

1 Like

Btw I've just run some tests locally and things seem to work!

Thanks Felipe!

I just addressed the feedback, pushed the changes and promoted the PRs from Draft to Ready for review.

I'll try run some tests just in case, but I think they're good for a pass now. Let me know if you find anything else.

Right after this I'll be back at looking how those three abstraction modes we discussed could work, but with after this changes I think we can discuss them independently and un-break the default debug mode for single threaded Embedded consumers.

Cheers.

Awesome! I've just kick-started CI runs

1 Like

I've merged the PRs!

One thing we may run into: I will soon need to annotate some of the debug-related variables with addresspace(1), so that they become symbols in wasm targets (and therefore visible to LLDB). But I'm not sure we can do that from swift code (looking mostly at EmbeddedPlatformSingleThreaded.swift)

I made this PR [Concurrency] Make EmbeddedPlatformSingleThreaded debugger-friendly to annotate some of the variables in that platform in a way that the debugger can read.

LLDB consumes that information in this other PR: [lldb] Add support for swiftEmbeddedPlatformSingleThreaded

I still don't know how this will work for WASM, where symbols don't get an entry into the symbol table unless they are annotated with addrspace(1), which is not available from swift. Investigating...

I see your point. I think it would be reasonable to move the debug variable’s definition to C if that helps with the Wasm annotations. The single-threaded PAL already imports the C declarations in EmbeddedPlatform.h, and the Darwin PAL defines the storage-kind variable in C, so this seems like a reasonable extension. Would a small C companion file for the debug metadata help here?

I think that could help, for sure. I currently have a lot of small fixes for debugging tasks in general, but everything is blocked waiting for the rebranch work to finish, as I don't want to merge anything potentially disruptive right now. Once that's done, I'll start merging and explore that idea too

1 Like

Hi Felipe,

Sorry for the long message! Tried to add enough detail so this post can be read in isolation. Please consider the implementation an experiment: it's only there as a way to discuss the debugger <> PAL contract, run tests, and try to poke holes or find rough edges.

Following up on the three lookup mechanisms we discussed, I now have an experimental Swift/LLDB implementation tested with real Embedded Swift tasks on a Raspberry Pi Pico 2’s two Cortex-M33 cores. Links at the end of the message.

This builds on the platform-deferred storage mechanism: the runtime delegates the storage description to the linked PAL, which selects how LLDB locates the current task.

Execution Contexts
All three mechanisms operate on an execution context, a platform entity whose current-task state we want to inspect: an OS or RTOS thread on a threaded system, or a hardware thread, such as a CPU core or hart, on a bare-metal system. An explicit (index, context_kind) pair is supplied by the platform debugger integration:

  • context_kind identifies the namespace: SOFTWARE_THREAD for an OS/RTOS thread, or HARDWARE_THREAD for a logical CPU or hart.
  • index is the platform-defined lookup key within that namespace. It identifies the execution context, not a Swift task.

For example, (1, HARDWARE_THREAD) could identify CPU 1, while (41, SOFTWARE_THREAD) could identify RTOS thread 41. For the helper-based mechanisms, the key could also be a thread-control-block pointer encoded as uintptr_t, provided the platform and debugger agree on its meaning and lifetime.

Lookup Mechanisms

There are three mechanisms, reflecting different storage guarantees a platform might provide:

Mechanism Contract Scenario
Indexed storage Publish one process-wide table through a base address, count, stride, and expected context kind. LLDB computes the selected context’s slot address using the index and reads its current task pointer without executing target code. Useful on platforms with a small, fixed number of execution contexts, like CPU cores.
Slot-address helper A helper is called with (index, context_kind) to obtain a stable task-pointer slot address. That address is cached for the execution-context lifetime, and its contents are read on subsequent queries. Useful for abstracting access to TLS when threads are recognized but their storage layout isn't.
Current-task helper A helper is called with (index, context_kind) on every query, and LLDB uses the returned task pointer. No stable storage address is required. Useful as a catch-all for any other platform whose storage model doesn't fit the other optimized cases.

Indexed storage
The first case is a small, known set of execution contexts, typically CPU cores on a bare-metal system. The platform maintains one process-wide table with a current-task pointer for each context and publishes its base address, count, and byte stride. LLDB computes the selected context’s slot address and reads it without executing target code.

Each entry contains an AsyncTask *, not a numeric task ID. The slot address remains fixed while its contents change as different tasks execute there. NULL means the context is idle; otherwise, LLDB obtains the task ID from the pointed-to task object.

The platform’s scheduler/runtime integration maintains these entries whenever the active Swift task changes, including nested job execution and restoration. A plain global array is sufficient.

Slot-address helper
The second case is a platform with threads and some native or emulated TLS mechanism that LLDB does not understand. An RTOS with its own threading implementation is a motivating example.

The platform provides a helper that takes (index, context_kind) and returns the address of that execution context’s current-task slot. LLDB calls it on a cache miss, caches the address, and reads the slot’s changing contents on subsequent queries.

This requires an explicit stability guarantee: the slot must remain at the same address and belong to the same execution context throughout its lifetime. Changing the task stored there or moving the thread between CPUs is fine. Relocating the slot is not. The cache tracks context lifetime and identity, including numeric ID reuse, but cannot reliably detect a slot that silently moves while its context remains alive.

A NULL helper return means unavailable and is not cached as valid. A valid slot containing NULL means idle.

I'm still working through how cache invalidation should work for this option. A reused thread ID should not retain an old cached address. I have some ideas that I'm weighing, but before getting deeper into that, I wanted to prioritize getting feedback on the three approaches.

Current-task helper
The third case is a platform that cannot expose a fixed table or guarantee a stable slot, but can answer which task belongs to an execution context.

Its helper takes the same pair and returns the task pointer on every query. This accommodates unusual storage arrangements, including platforms without conventional TLS, and prioritizes enabling task-aware debugging over the cost of repeated inferior calls. It also works when TLS exists but its storage is unsuitable for caching. NULL means idle; the prototype reserves UINTPTR_MAX for an unavailable lookup.

Calling Helpers
Both helpers execute in the target. Current LLDB schedules the call on the selected context’s backing LLDB thread, but this does not guarantee physical-CPU affinity for a software thread or preserve a whole-system snapshot.

The supplied pair is the authoritative lookup subject. Implementations must resolve it rather than substitute their current CPU or ambient TLS, unless the platform/debugger integration separately guarantees that they correspond. Calls require explicit opt-in, and helpers must not allocate, block, or acquire locks. The procedural fallback is flexible about storage, not about call safety.

Experimental Results
The hardware fixture uses CPicoSDK’s multicore scheduler and real Swift tasks. It moves a task from core 0 to core 1 and back, runs a second task, and checks repeated queries and idle contexts. All three mechanisms passed.

At a source breakpoint immediately after await Task.yield(), the same argument-free command produced these results:

# Unmodified debugger
(lldb) language swift task info
error: could not find the task address

# Patched debugger
(lldb) language swift task info
(UnsafeCurrentTask) current_task = id:1 flags:running {
  address = 0x20015980
  id = 1
  enqueuePriority = .medium
  parent = nil
  children = {}
}

The stop at Tasks.swift:24 ran on core 1; after another yield, line 26 ran on core 0. Both identified the same task ID and address. With task-as-thread presentation enabled, both appeared as Task 1 with the same synthetic thread identity.

Indexed lookup executed no helpers and preserved registers and both heartbeat counters during inspection. The slot helper executed once per execution context, which here means once per core. The value helper executed on every query. Helper calls restored the calling integer registers, but the peer core advanced with this OpenOCD setup.

The LLDB test run passed 129 variants: 60 indexed-storage variants, 48 helper variants, and 21 existing GDB-remote tests. Software-context behavior is covered by mock remote-server tests, not an actual RTOS integration. ARM32 logical async unwinding and complete task enumeration remain outside this experiment’s scope.

The experimental implementation is on the Swift draft PR and LLDB draft PR. And there's also a gist with a use example, very high level.

Would love to get some thoughts. I hope this is heading in a direction that could eventually land. I think the value in this approach is generalizing the problem while also creating some optimization opportunities. Looking forward to feedback and ideas on how this could be evaluated further.

1 Like

Let's take a step here for a second and imagine a world where there is no concurrency on Swift.

How would a program in these platforms be debugged? If you just launch a program that does not use the concurrency library in these systems, what does the LLDB Thread abstraction map to (e.g. if you break inside main and do thread list)? Is it just showing threads == CPUs?

I ask this question because I think the LLDB PR seems to be breaking a level of abstraction in LLDB, and this can be seen by the fact that the proposed code changes introduce swift-only abstractions inside the more general LLDB code. In particular, it is not clear to me why the mapping you are doing cannot be done inside the plugin OperatingSystemSwiftTasks, instead of at the Process level.

I've just realized that Github was hiding the diff on the SwiftLanguageRuntime file (:person_facepalming:), so I see that you are indeed essentially adding new "TaskFinders".

A few extra points:

  1. Is there any reason changes to {Process,Thread}GDBRemote.{h,cpp} were needed?
  2. The bits inside lldb/source/Plugins/ABI/ARM/ABISysV_arm.cpp are not too familiar for me, but it sounds like you are fixing an existing bug? Definitely deserves its own upstream llvm patch.
  3. You exposed a ExecutionContextIndex struct inside Thread.h, but that is not used anywhere in generic LLDB code. I understand you are trying to generalize the concept of TLS to something else, one way to do this would be to change how the lldb-server, debugserver, etc respond to the jExtendedThreadInfo packet. See ThreadGDBRemote::FetchThreadExtendedInfo, ProcessGDBRemote::GetExtendedInfoForThread. Maybe then things would flow a bit better?

These are just my initial reactions, I'd like to look at this again soon. I feel like we're missing a level of abstraction here, but don't quite know what that is. I don't know how much AI assisted code there is here, but if you have access to those tools, it would make things a lot easier to understand if the LLDB PR could be organised in logical independent commits. These tools are pretty good at breakpoint commits apart.

Hello Felipe, thanks for going through the post!

I don't know how much AI assisted code there is here, but if you have access to those tools, it would make things a lot easier to understand if the LLDB PR could be organised in logical independent commits. These tools are pretty good at breakpoint commits apart.

Fair, I can split the LLDB changes into logical commits. Just to reinforce the intent: the implementation is there to validate the idea and discuss the contract, not as a proposal for the final implementation. It's mostly AI-driven, pursuing the goal of getting the contract I provided to work. I'm sure there are many design aspects that need to be revisited, including where the different pieces belong in LLDB.

I feel like we're missing a level of abstraction here, but don't quite know what that is.

I think there are two slightly different questions here: what contract the PAL exposes, and where the integration belongs in LLDB.

For the contract, I'm defining two points:

  1. A platform lookup key, (index, context_kind), associated with an existing debugger-visible thread. This isn't intended to redefine LLDB's thread model, just to identify the context whose current Swift task we're trying to find.
  2. Three alternative ways to locate that context's current task pointer, selected through a storage-kind value exported by the PAL.

The goal is to let LLDB find the current AsyncTask * without needing to understand the platform's storage implementation. Once it has that pointer, the existing machinery can read the task information.

I think this could be abstracted further, but I'm not sure whether that would simplify the problem or just broaden it. That is separate from your point about where the integration should live, though. The prototype does introduce the lookup-key representation into generic Thread.h, and I don't think that placement is necessarily right.

Is the missing abstraction you're thinking of mainly at that LLDB boundary, or do you see something missing from the PAL contract too?

How would a program in these platforms be debugged? If you just launch a program that does not use the concurrency library in these systems, what does the LLDB Thread abstraction map to (e.g. if you break inside main and do thread list)? Is it just showing threads == CPUs?

In the Pico/OpenOCD setup I'm using, yes: the underlying LLDB threads correspond to the two Cortex-M33 cores. OpenOCD supplies that mapping independently of Swift concurrency.

That isn't necessarily the mapping on every embedded platform. With an RTOS-aware debug server, LLDB threads could instead represent RTOS threads. I'm assuming that baseline mapping already exists. The additional information needed here is how each existing thread maps to the lookup key understood by the PAL.

I ask this question because I think the LLDB PR seems to be breaking a level of abstraction in LLDB, and this can be seen by the fact that the proposed code changes introduce swift-only abstractions inside the more general LLDB code.

100%, that boundary needs another pass. The Process/Thread changes were experimental plumbing to expose the lookup key to the TaskFinders, not something I'm arguing must live there.

If there's alignment on the shape of the PAL contract, I can start looking into how to scope down and reorganize the implementation. I wanted to get feedback on that contract before spending too much time refining the implementation, but I understand the current organization makes that discussion harder.

What is a good way to have a deeper conversation about this? I can prepare a small presentation and bring it to a Working Group meeting, though I'm not sure which WG would apply.

Thanks Felipe!