[Pitch] `UncheckedString` (raw string support for Swift)

My experience as a French Canadian living in the US is that a truly shocking number of programs get encoding wrong. English speakers grossly underestimate how poorly programs deal with encodings on average because it never bites them. I hate making assumptions about people's backgrounds, but I think this might be yours. I have hundreds of emails in my work inbox calling me Félix, from multiple systems, and my success rate getting these very simple problems fixed is near 0. Many websites I sign up to tell me "[my] name is invalid". I regret to inform all of the linearly indexable UTF-8 String haters that it makes computers work better for me.

I understand the attractiveness of API similarity with String, but correctness-wise, UncheckedString is inferior to String. There are niceties that we can afford with String because we know the result will be correct. I'm not fundamentally against the idea that UncheckedString could support string interpolation, but I really think it warrants a real use case beyond it looking satisfying complete when you put the API side by side with String, and I think the arguments against it deserve more engagement than a shrug.

10 Likes

:-) To give some background, I've long been interested in character encodings and associated machinery, and I've actually got my own C++ string library that I wrote, with decent Unicode support and even the ability to deal with strings in random encodings without actually converting them to Unicode first (to the extent that it can even run Unicode regular expressions against e.g. a Shift-JIS string and come up with sensible matches).

Now, my name, obviously, doesn't contain accented characters and so I don't suffer from it in that sense, but I'm well aware that this is super frustrating (as are other internationalisation problems, like people asking for "first name", "last name", or making assumptions about address formats or telephone numbering schemes). This problem actually does exist for native English speakers too — we do have some accents, most commonly the diaeresis in names like Zoë or Chloë or words like naïve, but there are others (exposé, résumé, crème brûlée…) — but as you say it's less widespread.

Indeed, part of my motivation for this is that I'd like to fix CommandLine. Right now, it assumes that argv[] contains valid UTF-8 string data (or, on Windows, valid UTF-16), and will just do the wrong thing if it does not. This is not even guaranteed to be true on Apple platforms, let alone Linux or FreeBSD where you might have a terminal running in one encoding, looking at a filesystem where multiple users have saved files in multiple different encodings (even in the same directory), or where some mischievous person has saved files with names that just contain random bytes.

In addition, the way things are at the moment it's difficult to write a Swift program that takes e.g. a FilePath as a command line argument; swift-argument-parser doesn't support that type (though it's been a while — maybe that's been fixed since I last looked?), and even if it did, it starts with the strings from CommandLine…

Most people should be using String, and if this proposal is accepted, most people should continue to use String.

4 Likes

This is why I'm proposing that "Ren\x{e9} Descartes" be the debugDescription, not the description. In fact, UncheckedString in my view should not have a description. I think that's where the boundary lies — it's perfectly fine to have a debugDescription that prints approximately what you would have to write in your Swift program (including the quote characters). It's not fine to have a description because that implies you can successfully convert from the UncheckedString's unknown encoding to UTF-8, which you can't because you don't know how it's encoded. (Your program may, of course, know how it's encoded, but the standard library does not.)

4 Likes

In the absence of description, debugDescription is used as fallback for print, etc. So this would be an implementation-level distinction without a user-facing difference.

2 Likes

True, but I don't think it's the end of the world that converting to String with String(describing:) or using print() results in a quoted, escaped string. The point, surely, is that things that actually care can distinguish the fact that it isn't CustomStringConvertible.

1 Like

In that case, it sounds more like a string that is based on Unicode (or ASCII) but does not guarantee correctness, unlike String. That can be useful for file path components, HTTP, or other low level APIs. For a string with no encoding, I think I would still use Array or Data or something similar since it should not be interpreted as Unicode text by description and related APIs.

I saw you’re proposing a new conversion from UncheckedString to the various pointer types. It would be nice if we didn’t introduce any new specialized implicit conversions in the type checker. I think projecting a pointer from an UncheckedString should be an explicit operation, and we should move to that model for the existing String and Array conversions too.

2 Likes

That would make C interop extremely noisy IMO. The string literals are passed frequently in the Windows API, often as opaque handles (e.g. WC_LISTVIEW).

I suppose that an alternative would be to have an combinatorial explosion of overloads - for any function that takes wide string literals, provide a transparent thunk that does every combination of parameters as either the literal or the projection and maps to the projection. Would that be an acceptable overload set for type resolution?

As a concrete example we should then consider this:

int MessageBoxW(
  [in, optional] HWND    hWnd,
  [in, optional] LPCWSTR lpText,
  [in, optional] LPCWSTR lpCaption,
  [in]           UINT    uType
);

to be imported as:

  • (UnsafeRawPointer?, UnsafePointer<CWideChar>?, UnsafePointer<CWideChar>?, UINT) -> CInt
  • (UnsafeRawPointer?, UnsafeUncheckedString<UInt16>, UnsafePointer<CWideChar>?, UINT) -> CInt
  • (UnsafeRawPointer?, UnsafePointer<CWideChar>?, UnsafeUncheckedString<CWideChar>, UINT) -> CInt
  • (UnsafeRawPointer?, UnsafeUncheckedString<CWideChar>, UnsafeUnchedString<UInt16>, UInt) -> CInt

Where the last three are @_transparent wrappers that do the projection operation.

Don’t practically all of these call sites use the _T() macro in C++? You could define a @transparent func _T(_:) in a helper library. :slight_smile:

Nope; the vast majority are directly spelt with the wide string literal. _T is a different macro for a different purpose. The wide character literals are the de facto standard as UCS-2 is the natural base encoding (similar to Cocoa strings).

I thought the existing implicit pointer conversions were changed to only work for arguments to C* functions? Am I misremembering?

* and C++, Objective-C, and Objective-C++

You don’t need to use the L prefix?

For what it's worth I don't use wide character strings for Windows APIs, I just use ANSI strings, e.g. MessageBoxA, and never MessageBoxW. Works like a charm with Swift strings. The thing is that modern versions of Windows support UTF-8 encoding in ANSI strings.

2 Likes

This approach is endorsed by Microsoft, but only works if you set the active code page to UTF-8.

3 Likes

Not on the Swift side.

The ANSI Win32 APIs are documented as not supporting the Win32 bypass for long paths. Additionally, there are quirks with the file system with ANSI vs Unicode. As such, the core implementation must always use the W-variant.

How is this working right now? Do you have wrapper Swift APIs that take a String and pass its utf16view to WinAPI functions that take wchar_t *?

Stepping back, can we get a clear statement from @al45tair or the LSG on whether it is a goal or a *non-*goal of this proposal to stuff other encodings into an UncheckedString? As someone else pointed out, that’s really more like Python3’s bytes type, isn’t it? While UncheckedString might be a necessary intermediate representation to enable initializing a bytes from a string literal, that doesn’t mean it should be the currency type for passing around EBCDIC or UCS2 strings.

1 Like

It is emphatically not the case that it is intended to only hold Unicode or ASCII-based string data, though it is designed to make that common case convenient (hence the debugDescription behaviour). If you use it to store non-ASCII-based string data, then debugDescription will obviously produce "gobbledegook". So, for instance, assuming the existence of a CP500 EBCDIC codec:

let helloWorldUTF16: UncheckedString<UInt16> = "Hello, World! 😀"
let helloWorldEBCDIC = "Hello, World!".encode(as: CP500.self)

print(helloWorldUTF16.debugDescription)
print(helloWorldEBCDIC.debugDescription)

will output (assuming I have the encoding right)

"Hello, World! \x{d83d}\x{de00}"
"\x{c8}\x{85}\x{93}\x{93}\x{96}k@\x{e6}\x{96}\x{99}\x{93}\x{84}O"

The latter isn't particularly readable, but it also isn't going to feed unwanted control codes to your terminal or printer (for instance), and if you have a CP500 code table it's not too bad to work out what the string contains.

I will note that it is intended to be used for string data. Python programmers often use the equivalent type, bytes, to hold arbitrary binary data… I think it makes sense to use different types for that application for various reasons (e.g. the small string optimisation doesn't make much sense for raw byte buffers, there's no need to support NUL-termination, the debug presentation should be rather different and focus on the byte values rather than trying to present characters).

There are plenty of caveats here — for instance, old-fashioned GDI does not support UTF-8 narrow strings (a consequence being that anything built on top of it also won't), and third-party library support for this is likely extremely patchy.

Indeed, for which you need your program to have a manifest, which is something SwiftPM doesn't (yet) support.

There's also some question in my mind about whether or not running in this mode has character length limitations or overheads that don't exist with wide strings. The issue here is that, at least historically, the narrow APIs work by using MultiByteToWideString to convert the narrow string to a wide one, often in a fixed-size buffer on the stack. In cases where the buffer isn't fixed-size or it has fallback code to use a heap-allocated buffer if the fixed-size buffer isn't large enough, that adds even more overhead on top of the conversion.

For paths specifically, an important factor here is that it's unclear whether you can use long paths with the narrow APIs when in UTF-8 mode; it would make sense in principle that the opt-in path-length-limit removal that was added in Windows 10 version 1607 might work for UTF-8 mode, but the documentation doesn't make any such claim (and that setting requires both a manifest setting and a system-wide registry change too).

I think the signs are there that Microsoft is slowly pivoting to UTF-8, which is great. But I don't think we're quite at the point where it's a good choice in general for application developers. It likely works fine for a reasonable subset of programs, but it's definitely a work in progress.

I don't know about LSG, but I think it's a goal. There is no reason for UncheckedString to contain valid Unicode. It is a string of "characters" of unspecified encoding but with a well-defined code unit width. The \u{hh} escape, and the ability to initialise it with UTF-8 data from your source code is provided as a convenience for users… sometimes that might be what you want, and restricting it to ASCII in the source code seemed unnecessary.

In practice I'd expect most UncheckedString values to be ASCII-based, because that's the world in which we find ourselves.

I totally understand where you're coming from with this; those conversions are both horrible and somewhat dangerous (for lifetime related reasons), though at least they'll require the use of the unsafe keyword.

As @compnerd says though, without them, interfacing with C APIs becomes noisy and/or ends up requiring the generation of API stubs, with potentially enormous numbers of overloads, and this is an existing problem that we already have with String and Array. If we come up with a better solution for those, then that would apply here too, but I don't think we should try to do that as part of this proposal.

Very often Windows code ends up looking like this:

let result = "Are you sure?!".withCString(encodedAs: UTF16.self) { lpCaption in
  "You're about to delete all of the content from your computer.".withCString(encodedAs: UTF16.self) { lpText in
    MessageBoxW(hwndMain, lpText, lpCaption, MB_YESNO|MB_ICONWARNING)
  }
}

But this is inefficient and ugly; it'll convert strings at run time even when it doesn't need to and likely requires heap allocations to do so. That overhead doesn't go away if you wrap it into an overlay function.

In previous ownership discussions there was an example of + operator consuming its left operand. As consuming function allows to mutate instance, overall "a" + "b" + "c" become linear operation.

There were previous discussions about creating an ASCIIString type, but as far as I remember, they didn't lead to anything. For embedded systems and a number of other application scenarios, ASCIIString might be better suited than UncheckedString.

  • UncheckedString and UncheckedSubString both conform toRangeReplaceableCollection, so, unlike a bare Array<UInt8> or Data, they support in-place mutation (append, insert, remove(at:),replaceSubrange, reserveCapacity, +/+=, and so on) as well as construction from a fixed Collection.

FWIW it seems both Array and Data conform to RangeReplaceableCollection so they support all of those operations.

2 Likes

Indeed. It's not clear to me why we want a new type for this. We already have a few (arguably too many, depending on who you ask) names for "a contiguous RangeReplaceableCollection of homogeneous fixed-width data in an unspecified encoding with CoW value semantics."

The proposal appears to make no mention of Swift System's string types, which (together with Data) are the closest existing thing to this proposed addition that I'm aware of. Ticking down the list of other reasons given to introduce this type:

  • Arbitrary FixedWidthInteger code units.

This feels like the wrong protocol to target. Will it ever be meaningful to have strings of Int128 code units? I think every example given is either [U]Int8 or [U]Int16, right?

  • Efficient COW storage, with the small string optimization.

Small string optimizations is the main thing that Array<UIntN> won't get you today. How important do we think it is for the use cases in question? How much data do we have to inform this decision? Note the sleight of hand in putting "optimization" in the name of this design choice: it isn't always. If people take advantage of the RangeReplaceable conformance to write fiddly element-by-element operations, it will frequently be a pessimization. Data does it historically, and gets used as an argument against using Data instead of Array<UInt8>.

  • Compiler-supported string literal syntax, including a new hexadecimal escape sequence to allow the use of raw code units.

These are interesting! I'd like to hear from type checker folks about the potential impacts of the new literal support. Even if we don't add new types, better inits/literals on existing types might be pretty compelling.

  • encode()/decode() APIs to transform between UncheckedString and String.

These would be just as useful on Array<[U]IntN> or Data, right (more useful initially, actually, since there's lots of existing API that traffics in those types).

8 Likes