uchar.h — unicode
utilities header
The <uchar.h>
header provides support for the C11 and C23 Unicode utilities. The types and
functions provide means for working with data encoded as UTF-8, UTF-16, and
UTF-32.
The <uchar.h>
header defines the following types:
- char8_t
- An unsigned integer that can represent 8-bit characters encoded in
UTF-8.
- char16_t
- An unsigned integer that can represent 16-bit characters, generally a
single single UTF-16 code unit. A Unicode code point may be one or two
UTF-16 code units due to surrogate pairs.
- char32_t
- An unsigned integer that can represent 32-bit characters, generally a
single UTF-32 code unit.
- size_t
- An unsigned integer that represents the size of various objects. This can
hold the result of the
sizeof operator.
See also stddef.h(3HEAD).
- mbstate_t
- An object that holds the state for converting between character sequences
and wide characters (wchar_t,
char16_t, char32_t). See also,
wchar.h(3HEAD).
The
<uchar.h> header also
defines the following functions which are used to convert between
char8_t, char16_t, and
char32_t sequences and other character sequences. The
functions that end in
_l are extensions and
not part of the C standard. They take an arbitray locale to operate on
rather than using the current locale.
- c8rtomb(3C),
cr8tomb_l(3C)
- Convert char8_t sequences to multi-byte character
sequences.
- c16rtomb(3C),
cr16tomb_l(3C)
- Convert char16_t sequences to multi-byte character
sequences.
- c32rtomb(3C),
cr32tomb_l(3C)
- Convert char32_t sequences to multi-byte character
sequences.
- mbrtoc8(3C),
mbrtoc8_l(3C)
- Convert multi-byte character sequences to char8_t
sequences.
- mbrtoc16(3C),
mbrtoc16_l(3C)
- Convert multi-byte character sequences to char16_t
sequences.
- mbrtoc32(3C),
mbrtoc16_l(3C)
- Convert multi-byte character sequences to char32_t
sequences.