[Pitch] `UncheckedString` (raw string support for Swift)

That’s great for developers writing predominantly in English. What about developers writing in languages that make extensive use of high-ANSI characters? They can type these characters directly but will be forced to switch to \{xdd} notation because the Swift compiler doesn’t know what codepage will be in effect at runtime.

Not to mention code that interfaces with the whole host of non-ASCII-compatible encodings like Shift-JIS, Big5, EBCDIC, etc. And if we wanted to accommodate these use cases, the compiler needs more than a character width hint to tell it which encoding to target.

I fear this proposal reintroduces a whole host of concerns that Swift intentionally avoided by making the opinionated decision to be UTF-8 everywhere.

3 Likes

Here’s an idea for a completely different approach: a macro package that lets users type their desired string in native UTF-8, and have it transformed to an InlineArray at compile time.

let caption = #W(codePage: .windows1252, "Are you sure?!")
let text = #W(codePage: .windows1252, "You're about to delete all of the content from your computer.")
// Could also be `#W(unicode: .utf16, …)`, `#W(encoding: .ebcdic)` etc

let result = MessageBoxW(hwndMain, text, caption, MB_YESNO|MB_ICONWARNING)
4 Likes

That's a bad idea. If you need a non-ASCII character in your Windows program, you should be using the wide string APIs, and wide strings. It isn't just the compiler that doesn't know which ANSI code page is selected — the programmer doesn't either, because it's up to the user which ANSI code page is active.

Only if you expect to be able to write Unicode characters in the source and have them convert to the encoding of your choice. (Also, Shift-JIS and Big5 are both ASCII-based; certainly in the initial shift state, ASCII characters work just fine in those encodings, perhaps modulo the \ character which OTOH I think Shift-JIS might turn into a Yen symbol?)

I can't honestly say I'm a fan of that idea; it introduces an unnecessary macro, which will slow compilation and forces Windows developers to put ceremony around all of their strings (OK, it's better than the with pyramid of doom, but still). Also, how does this allow someone to specify a character that doesn't exist in Unicode? How is it better than being able to write

let result = MessageBoxW(hwndMain, "Are you sure?!",
    "You're about to delete all of the content from your computer.",
    MB_YESNO|MB_ICONWARNING)

and just have the compiler figure out that you needed wide string constants?

Are we talking about FilePath here? Because I think of that as a path type rather than a string type.

Agreed (I think there was some discussion about that above).

That, NUL termination (without keeping an extra NUL at the end by hand); compile-time literals; the ability to distinguish between an array of numbers or bytes and something that is some-kind-of-string; immortal values. I mean, maybe you're right and we don't need a new type — though I'd say from some of the comments above that it's not obvious that that's the case.

I think there is an argument that there really should be three different types of thing here:

  1. Arrays, e.g. Array<UInt16>. These hold a list of integers. If you were writing one, you'd write something like [1, 2, 3], and you'd expect e.g. debugDescription to generate something similar.
  2. Byte buffers, e.g. Data. These represent a chunk of memory. You might imagine a debugDescription something like:
    0000  00 01 02 03 04 05 06 07 08 09 0a 0b 0c 0d 0e 0f  ................
    0010  41 42 43 44 45 46 47 48 49 4a 4b 4c 4d 4e 4f 50  ABCDEFGHIJKLMNOP
    
  3. Strings. These represent text, and you'd write them like "Hello, World!" and expect debugDescription to format them in that manner. We might also provide affordances for interoperation with other languages for string types (like NUL termination), which we wouldn't want to do for arrays or byte buffers.

But aside from differences in syntax and serialization, the set of operations defined on each doesn't necessarily match (though there is clearly a huge amount of overlap). For example, calculating the sum of some character data, or uppercasing a list of numbers aren't operations that make a great deal of sense — what does it mean to uppercase 17, or to add "P" and "Q", after all?

I totally accept you could choose to unify them. C does, and Python unifies (byte-)strings and byte buffers with its bytes type too. So there's precedent for that. Personally, I'm not sure that's the best choice as it creates confusion about what the object you have is intended to represent.

For sure, I'd totally agree with that.

1 Like

Strongly agree, but it also does not generally make sense to sum bytes, nor can you uppercase bytes without an encoding. A “String” with an unknown encoding is in some ways much more like bytes than it is like a string.

5 Likes

:-) Yes, except even a string whose encoding is "unknown" usually doesn't have a totally unknown encoding. It might be reasonable to provide some affordances for text that is known somehow to be ASCII (or, for that matter EBCDIC) based, but whose encoding isn't otherwise known. I'm not proposing that here, but I can imagine wanting things like uppercaseASCII, lowercaseASCII, asciiCaseInsensitiveCompare(_:), as well as perhaps numeric conversions. IMO those would be wrong on Array<UInt8> and somewhat suspicious on Data.

Examples of the kinds of places that might be useful are HTTP or MIME headers or PDF parsing. There might also be some uses for these with debug information (which often contains strings, but usually doesn't specify what encoding they're in).

Other than the ExpressibleAsStringLiteral conformance, how is this substantially different from just Array<Element>? I realize they're not semantically the same, but is that enough to overcome the cost of introducing a whole new type?

2 Likes

At the risk of repeating myself:

  • The ability to distinguish in the type system between a string-like thing and an array of integers or a block of bytes, which enables different treatment of the data in certain cases (like serialization or debugging).
  • An efficient way to get a NUL-terminated C-style string pointer that doesn't require you to remember to put a NUL at the end of your array.
  • The small string optimisation (which avoids an extra heap allocation and subsequent indirection for small strings).
  • The ability to have immortal string constants (not just literals, but you can create these dynamically at runtime, if, for instance, you have some strings you know will never be deallocated but that you don't necessarily have at compile time as constants).
  • A more suitable debugDescription that knows that it's (likely ASCII-based) string data.

Sure, it's similar (indeed, the dynamic storage is actually backed by Array<Element>).

This makes it sound like there is a huge cost to doing that, which I don't think is true. Now, I do think we should have a high bar for putting things into the standard library, but I also think we should be open to adding new types where it makes sense to do so. Part of the point of having a type-safe language is that different kinds of data should be represented by different types, so that we don't mix them up. Using Array<UInt8> or Data to hold strings that might not be UTF-8 creates the opportunity for type confusion, is un-ergonomic in various ways and inefficient in others.

So the real question is:

  • Can we achieve a reasonable result using existing types, and what are the pros and cons of doing that versus adding a new type?

I can believe the answer to the question might be "yes", but if the starting point is Array<Element> or Data, then I think we need to be honest about the compromises that will create (or talk about how we're going to solve them).

Committing to Unicode (just without the validation guarantee) makes this type a lot more meaningful in my eyes:

  • The debugDescription is no longer guessing; errors in the output are errors in the data, not just the result of a wrong encoding guess.
  • One can add members (methods and properties) to this type that meaningfully treat it as a string. Otherwise, members could not do more than on an Array except when the member contains an encoding in its name or is passed an encoding parameter.
  • The encoding is no longer out of band (as is the case for Array or Data)
  • The conversions between String and UncheckedString could be well-defined, using init? in one direction and init in the other.

As currently proposed, it’s too easy to type a non-ASCII character in the string literal where it looks good in code but gets interpreted by a different encoding at runtime. This is a common issue in other languages and I see its effects every day. Swift’s String is better here and using the same convenient literal syntax for an “unchecked” API feels wrong.


That would be: swift-system/Sources/System/SystemString.swift at main · apple/swift-system · GitHub

I think a type like SystemString together with a narrower UncheckedString could be a good solution.

Part of the point of this type is that it is suitable for holding non-Unicode string data, including string data whose encoding isn't known or isn't yet known, so I don't think that's a good option.

I don't think we should have an "unchecked Unicode" string (that, it seems to me, invites the problems @fclout worries about), and the other thing that might be interesting, namely a string that knows its encoding (and maybe checks it) is also a different thing, as well as being potentially rather complicated to design and implement.

That might be an argument in favour of banning \u{hh} or non-ASCII characters from the literal syntax for UncheckedString and only permitting explicit \x{hh} escapes. The counter-argument is that for a string whose encoding is unspecified, the compiler doesn't know that it isn't UTF-8, UTF-16 or UCS-4, and if the programmer writes a character that requires Unicode, it's reasonable (and convenient) for the compiler to generate the appropriately encoded Unicode at that point.

This is particularly important for usability of wide string literals on Windows, I might add, where the developer knows that the encoding happens to be UTF-16, so writing Unicode characters is definitely safe. On the other hand, Windows doesn't enforce valid Unicode on its UCS-2/UTF-16-ish strings, so you probably don't want a checked UTF-16 string type, at least not for data that you get from Windows or its filesystem.

@scanon and I were talking about SystemString yesterday; it could certainly, with some tweaks, solve some of the problems I care about here (e.g. it might be suitable for command line arguments and environment variables, and with some additional work might also work for Windows' wide character API). It would still leave us with no appropriate type for 8-bit wide strings on Windows, though that's a fairly minor problem overall since we try to avoid those for various reasons; and it doesn't help with things like HTTP/MIME headers or parsing things like PDF — though maybe there are other solutions for those.

We do have UTF8Span which can be initialized without validating Unicode correctness, but as the name suggests it doesn't own its storage. I agree that I don't think introducing additional flavors of that would be helpful.

I wonder if this pitch should be divided into multiple proposals. Thinking about my earlier discussion, what I really want for my use cases is:

  1. A string literal that lets me express any byte in the range 0-255, because compiler performance for large array literals is prohibitive.
  2. An ExpressibleBy* protocol for that literal that is initialized by something like an UnsafeBufferPointer<UInt8> or an immortal Span<UInt8>, so that I can access that raw data with minimal overhead. (And there could potentially be an associated type for the element type to allow other things than UInt8.)

Everything else on top of that—the dynamic/small representations and other operations—are extras. But for what I'm trying to build, those two building blocks would be sufficient to build other APIs on top of it (and indeed, I would prefer to build my own type on top of that literal instead of relying on the standard library to provide something for me without overhead I don't want), and that would work equally well for the other types like UncheckedString.

3 Likes

Since the beginning, Swift has been highly opinionated that such data is not a string. Why all of a sudden is it a good idea to go back on that decade of precedent?

4 Likes

I'd draw a subtle distinction here; Swift has been opinionated that it is not a String, but that doesn't mean it isn't "string data" in a general sense. We aren't "going back" on anything by deciding that, in fact, it makes sense to have some string-like type that is able to hold string data that does not fit in a String.

Put another way, when it was decided that String should be a checked Unicode sequence, the reason wasn't that we wanted to make it especially awkward to process non-Unicode strings or for that matter, strings whose underlying encoding wasn't what we'd chosen for String (originally UTF-16, and these days UTF-8). It was that we wanted to make it hard to make the kinds of mistakes we'd seen developers make when manipulating Unicode strings in other languages.

2 Likes

The perfect is the enemy of the good.

And, if I may pontificate: we made a perfect[1] Unicode string type, but in a lot of ways it just isn't very good.

I realize it's an ABI and API nightmare, but I wish we could relax some of the constraints in String itself. Let me create a string with invalid Unicode with an initializer like init(unchecked:). Let it set a sticky bit in the String header that says "not validated" which propagates to other strings during concatenation/etc. Let me call try validate() manually.

Since we probably can't do that, having an additional string type with unchecked characteristics is the next best option, I suppose. But that makes me wary because the standard library already has an ever-growing number of string-shaped types. There are already a dozen ways in the standard library to hold a sequence of characters, mostly of Unicode provenance.

Is there a solution we could pursue that simplifies the landscape?


  1. YMMV. ↩︎

3 Likes

I feel like ExpressibleByBytesLiteral (with conformances for Data, [UInt8], etc.) + b"L\x{e1}tin 1\0" solves most of the problems this is trying to? The exception maybe being UTF-16 data, but I don't know how you'd manage that given you don't know the encoding or the endianness.

Indeed, we could extract the ExpressibleByUncheckedStringLiteral part of this proposal and then implement conformances on other things… that solves some of the problem, though as I've pointed out a few times already, if you're using Data or [UInt8] to hold the bytes of a string, the type system doesn't know that it's some kind of string. Knowing that is useful for serialization and debugging.

Additionally, Data and [UInt8] don't have the small string optimization, and I haven't looked to see how we might support constant literal storage for [UInt8] either (i.e. we might end up having to copy the data).

Nevertheless, it might still make sense to split the literal support out into its own proposal — the controversial part here seems to be having the UncheckedString type itself.

As regards this point, Unicode in the source code, whether directly embedded or using the \u{hh} escape, ends up as appropriate-sized Unicode in the literal constant, with endianness matching the system we're targeting. That does mean we don't support other-endian literals, though honestly I don't think that's a big problem. As regards not knowing the encoding, I don't see any reason why we shouldn't allow Unicode like this — it's up to the user to write appropriately encoded data for their use-case, and that might very well be Unicode or Unicode-like (e.g. WTF-8).

But on the other hand, for the use cases I'm thinking of, it's very much not a string of any kind. The data I want to pack is arbitrary binary data that doesn't correspond to any string encoding.

FWIW, I think one could reasonably argue that my problem could possibly be solved by hacking in special cases to the type checker. I'm not sure how viable that actually is (at the end of the day, if I have an array literal with 10,000 elements, it has to verify that they're all valid UInt8s). The optimizer is also really bad today at knowing when to outline the data vs. creating painful "allocate and append each element in sequence" codegen (and because it's done by the optimizer, you never get the efficient representation in debug builds).

I would hate to start depending on hypothetical type checker and codegen improvements, start generating these large arrays, and then have my entire ecosystem brought to its knees by a regression in a future version of the compiler. The appeal of a binary literal is that it eliminates a whole class of type-checker and codegen unpredictability for this use case, and none of it relies on having a "string-ish" type provided by the stdlib.

This is orthogonal to the problem of representing raw binary data in source code. If a type that conforms to the literal protocol wants to provide that optimization then it can, but there are also types that want to guarantee that the data just ends up in a stable address in the binary.

I think there's a very clear case for the literal syntax and protocol parts of the proposal to be split into their own proposal.

I’m not very familiar in practice with the problems that UncheckedString is trying to solve, but reading the proposal and the discussions in this thread gives me the impression that the overall complexity of the compiler changes (such as the entire ExpressibleBy- set of protocols) isn’t worth the benefits these changes provide. Ultimately, a set of bytes remains just a set of bytes.

It seems to me that all of this is more of a convenience API than a fundamental type like String. UncheckedString, by its nature, is closer to Array<UInt(8/16/32)> than to String.

In conclusion, it makes sense to add such an API to a library like SwiftSystem, Foundation, SwiftCollections or a separate tiny package.

Small object optimization: Data did have it until this year, and its performance impact was one of the most commonly-heard complaints about Data vis-a-vis Array. It has been removed on non-ABI-stable platforms (Streamlined New Data ABI by jmschonfeld · Pull Request #1706 · swiftlang/swift-foundation · GitHub). The small representation case is more compelling for string data than for most other types, but it adds very real overhead to some operations.

Constant literal storage: there is no way to make [UInt8] support this on ABI-stable platforms. On non-ABI stable platforms something would be possible, but you are almost surely better off using either UniqueArray or Span.

4 Likes