This is a new proposal to add an UncheckedString<Element> type that can be used to hold strings that are not necessarily valid UTF-8, including strings whose encoding isn't known, and strings with wide characters.
I still need to give this a full read but this is a feature I desperately want, and I'm glad the literal \x{...} syntax is supported here.
My use case is in swift-protobuf; I'm encoding potentially large schemas (how a message is laid out) currently in StaticStrings, because array literals are currently intractable: the type checker drags having to check each element of the array and the codegen is unpredictable/suboptimal. String literals solve all of those problems, with the caveat that I had to invent a 7-bit safe encoding to pack things tightly but avoid non-UTF-8 code units.
StaticString's utf8Start guarantees that I can get a pointer to that read-only data and work with it efficiently. I see the implementation section discussing various storage representations like the small-string optimization, so I'm not sure how those interact with literal data. Is a literal always treated as an immortal string, even if it would fit in the small representation? I can imagine this being frequently used for use cases similar to mine—embedded binary data in Swift source—so representational predictability is something that's needed. If UncheckedString<UInt8>("smol") is represented differently UncheckedString<UInt8>("something longer..."), then perhaps there could be an initializer that forces the immortal representation regardless of length?
Ultimately, can we provide a strong guarantee, independent of compilation mode (optimized or unoptimized) that an UncheckedString initialized by a literal must always provide efficient access to its underlying elements? (Ideally, nothing more complex than indexing into a pointer in read-only data.)
Similarly, it's not mentioned in the proposal, but does this type support section control attributes?
@section("__foo")
let mySpecialData: UncheckedString<UInt8> = "blah blah blah"
Yes. More specifically, the compiler uses _ExpressibleByBuiltinUncheckedStringLiteral, which uses the init(immortalString:, nulTerminated:) initializer.
Unlike with String, I decided to expose the immortalString initializer, so you can indeed create your own from data. I can imagine that might be a controversial decision, mind, so it might not make it into the final version.
Yes. And indexing the other forms is also efficient. Even the small string form is relatively fast (there's even an optimisation for 8-bit strings that lets you get a pointer directly to the small string data; we can't do that for wide strings, because of alignment issues, but for 8-bit strings we can).
I imagine you're asking about the literal data itself rather than the variable? The problem here is that the variable is the thing with the attribute on it — the literal constant doesn't have that attribute.
This is the same as for ordinary String literals, I might add.
Performance-wise, the current implementation I have of UncheckedString is at least competitive with String, and often significantly faster (though I'm not pitching this as a way to write faster Swift programs; I'm specifically after byte strings, for things like command line arguments, filenames, and environment variables; and 16-bit string constants for Windows work).
Right, I saw the initializer that takes a pointer as an argument, but for data embedded into the source code itself, the string literal case was my main concern.
However, the immortalString: initializer does open up the door to using the same type for schemas allocated dynamically at runtime and managed in a pool, so that's a very nice addition that would unify everything I want to do.
Fair point. I keep thinking about this in C terms where you just declare a const char * and so the only reasonable target of the attribute is the data, not the variable that holds the pointer to it. I guess we need an extension of @section's syntax or some other way to express a section-bound string literal here.
Actually, thinking about it and looking at the code, this is not true (it'd make the small string optimisation a bit pointless, after all, since you'd only be able to construct them dynamically). The string variable itself will indeed use the small representation in that case (the immortalString: initializer will return strings that are .empty, .small or .immortal; it never returns a .dynamic string).
But access to the bytes (for 8-bit strings) is still efficient. For wide small strings, they need to be copied out into local storage first; since they are small, this isn't a huge overhead, but it's worth mentioning.
This would certainly be useful for Windows. There are a number of places where UTF16 constants are required (e.g. #define CONSTANT L"constant") as well as places where we need UTF16 literals (e.g. GetModuleHandleW).
Overall, I understand why this is super important for Windows, but as a person whose name reliably reveals encoding bugs, I am not that enthusiastic.
CustomUncheckedStringConvertible with no notion of encoding will cause problems. I understand why you would create an UncheckedString from dubious bytes that you received at no fault of your own, but creating dubious bytes from a starting point you control is asking for trouble. There is intentionally no agreement between UncheckedStrings in your address space about the encoding that they each use, so generically, you never know if embedding an UncheckedString into another one has a consistent result.
I’d like to see what use cases people have for CustomUncheckedStringConvertible before we commit to it. It’s not very hard already to append dubious bytes to the dubious bytes collection, I don’t know that we need to bless the practice with extensible string interpolation.
In my mind, the only reasonable use of unchecked string interpolation is to put preexisting unchecked strings that came from the same source into each other. Creating an unchecked string from just about anything, which CustomUncheckedStringConvertible encourages, is problematic. This seems to be partially acknowledged by the fact that String isn’t CustomUncheckedStringConvertible.
Less wordy questions/comments, roughly in order of appearance in the proposal, below:
UncheckedString is not a great name because Swift is already using “unchecked” for memory-unsafe operations. If I had to bring my own paint to the bike shed, I’d pitch EncodedString. (Encoded how? Who knows!)
The proposal says that UncheckedString supports any FixedWidthInteger, but it only specifies how literals work when Element is UInt8, UInt16 or UInt32. What happens if I use "Félix" as UncheckedString<Int8>: do I have negative code units? If I use "Félix" as UncheckedString<UInt64>, do I just have absurdly large characters?
Am I understanding correctly that for UncheckedString literals, the compiler encodes the literal part to Unicode, and if you want a different encoding you use \x{hh} for the relevant parts?
Can I use a FixedWidthInteger type that is not from the standard library?
The proposal says you can use an arbitrarily long hex string in \x{hh}, as long as the element type permits it. What happens if you use a hex string too long: is it silently accepted and truncated, or do you have a compile-time error, or do you have a runtime failure?
What is the interaction between \x{hh} and signed Element types? Bit-cast to the element type?
Should your various span-closure-taking methods be borrowing properties instead?
Do you need special support to create UncheckedString<UInt8> from CChar pointers when you could just have an UncheckedString<CChar>? It seems you’ll need to use UncheckedString<CChar> to benefit from existing implicit conversions to pointer types anyway. CChar is Int8 on Darwin.
Would an UncheckedString<Int8>’s Codable representation encode the same way that Data does too?
Should you have an initializer that uses an OutputSpan-taking closure?
Technically, UncheckedString<Int8> is very reasonable: that is a C string. CChar is not guaranteed to be unsigned and in fact some platforms (e.g. BSDs) default to -fsigned-char. So, UncheckedString<Int8> is simply the proper spelling for a C string. However, as the signedness of char is an ABI contract and implementation defined, you shouldn't use that spelling (e.g. UncheckedString<UInt8> is the spelling for a C string on Windows).
That's absolutely true. It's Unchecked, so you, the developer, have to be careful.
That's totally fair. I wasn't proposing to provide any conformances to CustomUncheckedStringConvertible besides UncheckedString itself, and only then because interpolation is an efficient way to construct strings piecewise (if we don't have it, people will write "a" + "b" + "c" + someThing + "d" + someOtherThing without realising that that's O(n²) and therefore slow; persuading people to build their strings with += instead is an uphill struggle that has been lost repeatedly in other languages).
I can imagine that it might be reasonable to provide CustomUncheckedStringConvertible on e.g. integer types on the understanding that it generates ASCII (the Platform Steering Group generally ruled out EBCDIC support, after all), and things like web server implementations might want it for e.g. cookie values.
But yes, you're right to highlight this — UncheckedString is explicitly a bag of code units in no particular encoding.
I spent a long time thinking about the name. UnsafeString would be wrong because it isn't "unsafe" in the Swift sense. RawString is wrong because we have "raw string syntax" (which ignores escape sequences). I think EncodedString is wrong too because we don't know the encoding — I can imagine someone actually wanting that name for a type that does know the encoding. We already use the word "unchecked" in the continuation APIs, and the way we use it is pretty much the way it's used here — we're telling the user that they're on their own.
Yes to both. Also, some things may be slower if you pick a type other than UInt8, UInt16, UInt32, CChar (or on Windows CWideChar), and I think you will find that UInt64 string constants aren't implemented in the compiler either (at least in my current patch), so that will result in a compile-time error. Maybe we should have a new protocol that just encompasses the types we expect people to use here, rather than using FixedWidthInteger? (The implementation expects 8-bit, 16-bit and 32-bit FixedWidthIntegers specifically; some parts of it would work with larger units, but perhaps not all.)
I'm quite interested to hear what people think about that.
Your source code is Unicode. So for a literal, if you write Unicode data, then yes, the compiler will represent that as the appropriate form of Unicode according to the element width. If you want something else, you can indeed use \x{hh}. I'm not sure I emphasised this either, but \u{hh} syntax still works as well, and will expand to the appropriate Unicode data.
I don't think literals will work (the compiler needs to know the bit-width of the type, and right now to do that I'm checking against the [U]Int(8|16|32) definitions. Other things should work, but it'll likely be somewhat slower.
Again, maybe this tells us that the restriction should be something other than FixedWidthInteger.
Essentially, yes.
Interesting question. I hadn't considered using borrowing properties. The span-closure pattern is closer to what the existing String type provides, but maybe you're right. Interested to hear what people think about that too.
CChar is sometimesUInt8 instead, so I wanted to make sure that we interoperate. Also, I think the default should be UncheckedString<UInt8> — the fact C's char is signed is (for character data) all kinds of wrong… people just don't expect that. I don't know about you, but I've never seen a code chart drawn with 0x80 as the top row.
I think I covered Codable support in the proposal; the default is that it will encode itself as an array of its element type. If you have a custom encoder, then you can obviously have special support for it the way some things do for Data. Maybe I've misunderstood the question?
I'm not opposed; when would this be useful, do you think, given the current range of initializers?
Thanks for all of these — they're really good questions and I know I'm going to go away and think about them further.
I think I can deal with interpolating UncheckedString into UncheckedString, I'm not too into the extension points though. If we make interpolation extensible, we are giving developers permission to extend it. Since this can easily be mishandled, without a use case, I'd suggest we see if we can stick to interpolation working only with UncheckedString, and ExpressibleByUncheckedStringLiteral being underscored until there's some demand for it.
Well, sure, but it can easily be more complex than "it's Unicode because the source bytes are Unicode". If a UTF-8 source file contains the bytes 65cc81 (e\u{0301}, e + combining acute accent), will the UncheckedString contain 65cc81 or c3a9 (é)?
I haven't seen a mention of \u{hh} in the proposal, so I think that should be clarified, yes (along with whether that interacts with normalization, if normalization applies at all to characters in source).
Yes, IMO that is the correct outcome.
I agree, but at the same time I only look at the numeric value of a char when it's used as an integer, and values outside of 0...127 nowadays typically have no textual meaning on their own. We've already established that \x{hh} bit-casts to match signedness, so I wonder how important it is that the canonical representation of a C string is UncheckedString<UInt8> rather than UncheckedString<CChar>. FWIW, I think a solid case could be made either way, but I think that's been under-explored.
I think I just got confused by the phrasing:
the same representation Array<Element> would use, and (for UInt8) the same representation Data uses.
Not being familiar with how Data encodes, I thought this said that Data and [UInt8] encode differently and UncheckedString<UInt8> aligns with Data, when in fact this is saying that Data, Array<UInt8> and UncheckedString<UInt8> all encode the same.
There is one on String and I think @glessard wants to extend them to most containers, so that seems indicated. If he's not in touch yet, you should ask his thoughts about how your proposal interacts with Span as he'll probably have more ideas than me.
The proposal itself looks good and it's definitely a +1, as someone who has and is still working on a IDN library which accepts any bytes and needs to perform some unicode operations such as checking for or converting to NFC, and performing Puny-code operations.
Two notes:
/// Calls the given closure with a buffer of `Element`s,
/// which are *not* necessarily NUL-terminated.
@inlinable
public func withCharacterData<R, Failure>(
_ body: (Span<Element>) throws(Failure) -> R
) throws(Failure) -> R
Sounds to me these kinds of with... functions can instead be some vars?
I don't know if there is an actual blocker but I'd assume it's possible to provide some variables instead. Similar to Requirements for IP address and port APIs - #32 by MahdiBM.
public var characterData: Span<Element> {
@_lifetime(borrow self)
get {
let pointer = UnsafeRawPointer(Builtin.addressOfBorrow(self.whatever))
let span = unsafe Span<UInt8>(_unsafeStart: pointer, byteCount: self.elementCount)
return unsafe _overrideLifetime(span, borrowing: self)
}
}
While It's not really a shortcoming of this proposal, I still would like to see considerations or at least awareness regarding the hopefully upcoming "bag of bytes" types.
Similar to this comment and the related to the discussion [Discussion] Bag of Bytes Types - #25 by MahdiBM.
Essentially, I'd like to see at least escape hatches for not having to copy memory around.
For example if you get some data from C, there should be some possibly unsafe way to pass that data to UncheckedString without copying the data.
Also this'll be helpful in other contexts like with SwiftNIO which passes data to you via its own ByteBuffer type.
This one should hopefully not be needed anymore in some future version of Swift where we have the bag of bytes types. It'll need to be enabled via NIO's ByteBuffer being able to freely transform itself to OurNewBagOfBytes (very possible IMO) and UncheckedString being able to accept and take care of that OurNewBagOfBytes without copies or such (Depends on the impl).
Looking at the proposal, I'd say a OurNewBagOfBytes should not be a big deal for UncheckedString as UncheckedString can just add an enum case that accepts it.
However, we'll need UncheckedString to not become frozen before then.
The downsides here are significant; there is no literal support, whatsoever
Overarching question: to what extent is the must-have feature here the raw string literal, with the rest of the pitch reflecting the issue that we don't have a single go-to currency type for a "bag of bytes" (or bag of UInt16s) to tie the literal to?
I raise this question because the pitch adds a lot of "stuff" to the standard library: a generic type, a protocol for that type, an ExpressibleBy... protocol, a ...Convertible protocol, and possibly another protocol for the subset of fixed-width integer element types, and probably more... But, meanwhile, the go-to thing that both the pitch itself explaining the major win over existing types and @allevato's comment in response jump straight to the point that the Big Thing here is the raw string literal.
Would it be more focused here to design a really good raw string literal for Swift, and then work on the ecosystem problem of nailing down a suitable "bag of bytes" type as a parallel effort? As it is, the conversation seems to be focusing on designing all the ancillary "stuff"—but maybe that's not the crux of what's to gain here?
My intent here was to make the interface approximate (as closely as possible) that of String. I think it's useful to be able to extend it. I can imagine it being reasonable to allow interpolation of Data or similar types into an UncheckedString, for instance, as well as potentially reasonable to support interpolation of numeric types to their ASCII equivalents. I'm not proposing adding that functionality as part of this proposal, mind.
65 cc 81, in both cases. We'd have to normalise to get c3 a9 and that would be wrong (IMO). If you want c3 a9, you can of course write \u{e9} or é rather than e\u{0301} or é.
It is mentioned:
I don't believe we normalise characters in string literals when they're read by the lexer. That would be undesirable, I think.
The point here is that your source code is UTF-8, and any Unicode things you write in a string literal that happens to be an UncheckedString will turn into the appropriate UTF-8/UTF-16/UCS-4 sequences in your UncheckedString literal. That applies whether you wrote them with \u{hh} or whether you wrote Unicode characters directly into your source code.
I think UncheckedString<UInt8> is a better choice for Swift. The problem with CChar is the same as the problem with char in C — you don't know if it's signed or not, and actually for character data, treating it as signed is a persistent source of bugs (c.f. using it to hold integer values, where it's often more useful for it to be signed). Stating up front that we're going to default to UncheckedString<UInt8>, and allowing that to interoperate with CChar pointers where necessary, makes sense from that perspective, I think.
Interesting question. I think it likely depends on whether you think that a bag-of-bytes type should have string operations and string literals, or whether you think that it should be a separate type entirely.
I wasn't thinking of UncheckedString as a container of random data, more as a container that holds something that is a string, but whose encoding we don't necessarily know — but in principle it's true that you could unify UncheckedString with some bag-of-bytes type (Python went down this road — its bytes type is basically just that and gets used to hold arbitrary data rather than just byte-strings). There are definitely pros and cons of unifying these; for instance, the description or debugDescription for a block of data might generate something like
0000000: 01 23 45 67 89 ab cd ef .#Eg.«Íï
0000008: 04 00 00 00 02 10 00 00 ........
while for a string it might make sense to do something like
"Ren\x{e9} Descartes"
There are other things to think about here too; do we want bag-of-bytes types to support C-style NUL-terminated strings via constructors? Should they keep their bytes NUL-terminated in memory just in case they're asked for a NUL-terminated string (a string implementation probably should do this, but I'd say a byte buffer likely shouldn't, not least because it would mean that a page-sized byte buffer would then run over into a second page because of the trailing NUL).
This is definitely a choice, and to some extent a philosophical one — what is a "string", and what is "data"? Is there a difference? Different languages have made different choices here, for various reasons.
To be clear, it's only adding all the "stuff" to match (roughly) what we have for String, so the new stuff isn't so much a new design as a clone of the existing design for a new kind of string type.
Love this effort, mentioned the need for new encodings before, so I hope this would be a good way to use them. I find Unicode to be a giant ill-considered hack, and think it was a mistake for Swift and other newer languages to standardize on it. I hope this will help build new encodings that lead us out of the current mess.
Yeah, I think one of the strengths of Swift's stdlib design is that we've taken the opportunities afforded by our modern era to say things like: "Integers are two's complement, floats are IEEE, and strings are Unicode." It's hugely clarifying and, I think, good for users.
In that scheme, a "string of unknown encoding" is a block of data. I think it makes great sense to allow users to provide, as input, "Ren\x{e9} Descartes" (and hence my point that the biggest "get" of this proposal is the raw string literal) just as it makes sense to allow users to convert arbitrary strings of known encoding so that the encoding is "opaque." It seems to me that many (or even most) operations you'd want to do with such a block of data that could be a string of some encoding become much more fluid with that one affordance.
But I think there is a conceptual boundary there we should not cross: it would muddy the waters for "Ren\x{e9} Descartes" to be the default output or description. By saying that the encoding is unknown, we are saying there is no sure "R" or "e" in the data: we can certainly make a "possibly ASCII" or "possibly Unicode" view of that data ergonomic to access, but there is a category difference between that approach and duplicating the string hierarchy for this type.
Yeah, no: I think we should rule this out from the outset as a non-goal. We're not displacing IEEE floats or ASCII-compatible encodings any time soon, and we should absolutely not be constructing parallel protocol hierarchies on that notion. Recall that the Space Shuttle was sized for railway gauges based on the wheel spacing of Roman war chariots. And this process is called Swift Evolution, not revolution.
As Steve wrote in the other thread, when (if) another viable universal encoding standard presents itself, our approach would be to provide new types that concretely model that other standard; it would not be to have existing types become wishy-washy. We have enough concrete use cases here with existing data to be designing to notional future use cases.
Sure, nowhere did I say that should be a goal, but rather something that people can already do with existing types that would further be enabled by this effort to allow alternate string encodings.
It is fascinating that you chose that example, as it is exactly my example of why those physical standards made some sense, but current digital ones do not. In this day and age when even your toaster is updating its firmware (updated with an actual toaster ), there is no reason at all to prematurely standardize. Software standards like Unicode come from an outdated pre-internet era when it was necessary to standardize a bunch of non-networked computers as much as possible to save costs.
That is not the world we mostly live in today: now all that matters are de facto standards, though hopefully they are still well-specified. I'll get off my soapbox now, no need to get this thread sidetracked by what encodings this will enable or whether the Unicode standard is worthwhile.
I agree with this philosophically, but find that boundary to be difficult to agree to unless we have implicit conversions with unchecked encoding specifications. Something like UncheckedUCS2String(“Ren\x{e9} Descartes”). The representational boundary is unfortunately a core need for dealing with UCS2/UTF16 strings on Windows, and sadly print debugging is here to stay.