C# example

How to convert a string to a byte array in C#

7 min read Updated Oct 2026 Runs in an isolated runtime
Quick answer

Use Encoding.UTF8.GetBytes(text) to turn a string into a byte[], and Encoding.UTF8.GetString(bytes) to turn it back. Both live in System.Text. Decode with the same encoding you encoded with: UTF-8 bytes read as Latin-1 turn "Café" into "Café". Don't call bytes.ToString() for the reverse; it prints System.Byte[].

A C# string is a sequence of UTF-16 chars, and a byte[] is just numbers. Getting from one to the other always goes through an encoding, a rule for which bytes stand for which characters. There is no cast, and no "natural" byte form of a string that every program agrees on. The Encoding class covers the common ones, and UTF-8 is the right default for files, network protocols, JSON and hashing. The rest of this page is about the times the default isn't enough: another encoding, the string's raw in-memory bytes, a byte order mark at the start of a file, and bytes that aren't valid text. Each example runs on this page: hit Run, then edit the code and run it again.

1Encoding.UTF8.GetBytes and GetStringRecommended

Encoding.UTF8.GetBytes encodes every character and returns a new array. Encoding.UTF8.GetString does the reverse, either for the whole array or for a slice given as a start index and a byte count. ASCII characters take one byte each in UTF-8, and anything else takes two to four, so the byte count is often not the string's Length. For a constant, C# 11 added u8 literals, which the compiler encodes for you into a ReadOnlySpan<byte>.

Program.cs

Output

Prints 12 chars -> 14 bytes: é and ö take two bytes each (195 169 and 195 182, or C3A9 and C3B6 in hex). Decoding gives back héllo, wörld and the comparison prints True. The slice prints héllo, because the first six bytes are five characters. Slice on character boundaries: a count that cuts a character in half decodes to a replacement character (section 5). Then System.Byte[], the type name, which is all ToString() on an array gives you. The u8 literal is 6 bytes and decodes to héllo. For a readable dump, Convert.ToHexString (.NET 5+) is the one-liner; see converting a byte array to a hex string.

2Other encodings: UTF-16, UTF-32, Latin-1 and ASCII

Every Encoding has the same GetBytes and GetString methods, so switching is one word. Use the one the other side expects: a file format, a protocol or a legacy system usually names it. The loop below encodes the same string with each built-in encoding and decodes it again, then decodes UTF-8 bytes with the wrong encodings to show what that looks like.

Program.cs

Output

The 7-character string is 10 bytes in UTF-8, 14 in UTF-16 (Encoding.Unicode, little-endian, and Encoding.BigEndianUnicode) and 28 in UTF-32. Those four round-trip to Café €5. The two single-byte encodings don't: Latin-1 has no € and prints Café ?5, and ASCII stops at 127 and prints Caf? ?5. Neither throws; the encoder writes a ? and moves on, so the loss is silent. The last three lines are the classic mojibake: UTF-8 bytes 436166C3A9 read as Latin-1 give Café, and read as ASCII give Caf??. Encoding.Latin1 is .NET 5+; Windows code pages such as 1252 need a provider (see the FAQ).

3The raw bytes, without choosing an encoding

Sometimes you don't care what the bytes mean. You want a byte form of the string to hash, cache or store, and to get the exact same string back later. The string already has one: its UTF-16 code units in memory. MemoryMarshal.AsBytes(text.AsSpan()) (.NET Core 2.1+) views them as bytes without copying, and MemoryMarshal.Cast<byte, char> goes back. Older code copies the chars with Buffer.BlockCopy, which gives the same bytes.

Program.cs

Output

Prints 480069002000AC20: two bytes per char, low byte first, so € (U+20AC) is AC20. That is 8 bytes, the BlockCopy copy matches (True), and on a little-endian machine, which covers x64 and ARM64 as .NET runs on them, it is byte for byte Encoding.Unicode (True, True). Back comes Hi €. The difference shows up with a broken string: a lone surrogate D800 survives the raw view (610000D86200), while Encoding.Unicode swaps it for U+FFFD (6100FDFF6200) and UTF-8 writes that as EFBFBD. So the raw bytes always round-trip, but they depend on byte order and are twice the size of UTF-8 for English text. Anything that leaves the process should use a named encoding.

4The UTF-8 BOM in bytes from a file

Some UTF-8 files start with a byte order mark, the three bytes EF BB BF. Notepad on older Windows and many Excel exports write it. File.ReadAllBytes hands you those bytes as they are, and GetString doesn't skip them: the BOM becomes an invisible U+FEFF character at the start of the string. Readers that know about it, File.ReadAllText and StreamReader, detect it and drop it.

Program.cs

Output

The file with the BOM is EFBBBF68656C6C6F and the default one is 68656C6C6F. Decoding the first gives 6 False U+FEFF: the string looks like hello when printed but has six characters and isn't equal to it, which breaks comparisons, CSV headers and JSON parsers. Skipping Encoding.UTF8.Preamble by hand gives True, File.ReadAllText gives True, and the StreamReader reads 5 characters. GetBytes never writes a BOM (5 bytes for hello), but passing Encoding.UTF8 to File.WriteAllText or a StreamWriter does. Use new UTF8Encoding(false), or leave the encoding out, to write files without one.

5Invalid bytes and bytes that arrive in chunks

Not every byte array is valid UTF-8. By default GetString doesn't throw: it replaces each bad sequence with U+FFFD, the replacement character, and carries on. That is the right call for display, and the wrong one when bad bytes mean corrupt data. For that, build a UTF8Encoding with throwOnInvalidBytes: true. A second way to get replacement characters is decoding a stream chunk by chunk, because a multi-byte character can straddle two chunks. A Decoder holds on to the unfinished bytes between calls.

Program.cs

Output

The lenient decode prints Hi�!�, which is 0048 0069 FFFD 0021 FFFD: one replacement for the stray FF and one for the unfinished C3. The strict encoding throws DecoderFallbackException: Unable to translate bytes [FF] at index 2 from specified code page to Unicode. Decoding the two halves of € separately prints ��, two replacements and no euro sign. The Decoder returns 0 chars for the first chunk, keeps its two bytes, and returns 1 char once the last byte arrives, so the result is €. You rarely need one by hand: StreamReader uses a Decoder internally, so reading text from a Stream through it handles the split for you.

6Which should you use?

MethodBytes per charRound-trips any stringBest for
Encoding.UTF8.GetBytes / GetString1 to 4Yes, for well-formed textAlmost everything: files, network, JSON, hashing
"text"u81 to 4Yes (constants only)Fixed UTF-8 bytes with no runtime encoding (C# 11+)
Encoding.Unicode / BigEndianUnicode2 or 4Yes, for well-formed textWindows APIs and formats that specify UTF-16
MemoryMarshal.AsBytes(s.AsSpan())2Yes, even lone surrogatesIn-process hashing or caching, no copy
Encoding.Latin1 / ASCII1No, other chars become ?Legacy protocols that require them
new UTF8Encoding(false, true)1 to 4Throws on bad bytesDecoding data where corruption must fail

Frequently asked questions

Should I use Encoding.Default or Encoding.UTF8?

Use Encoding.UTF8 and say what you mean. On .NET Core and .NET 5+, Encoding.Default is UTF-8 too (its WebName is utf-8), but on .NET Framework it is the machine's ANSI code page, so the same code produces different bytes on different machines. One difference on modern .NET: Encoding.UTF8.GetPreamble() returns the 3-byte BOM and Encoding.Default.GetPreamble() returns none, which matters when you pass them to a StreamWriter.

Why does ToString() on a byte array print System.Byte[]?

Arrays don't override ToString(), so you get the type name: new byte[] { 72, 105 }.ToString() is System.Byte[]. To turn the bytes into text, decode them with the encoding they were written in, for example Encoding.ASCII.GetString(bytes), which gives Hi. To show the numbers, use string.Join(" ", bytes) or Convert.ToHexString(bytes).

How do I store a byte array that isn't text, like a hash or an image, in a string?

Use Base64 or hex, never Encoding.UTF8.GetString. Arbitrary bytes are usually not valid UTF-8, and the invalid parts become U+FFFD: the bytes FF 00 80 decode and re-encode as EFBFBD00EFBFBD, so the original is gone. Convert.ToBase64String gives /wCA and Convert.FromBase64String brings back exactly FF0080. Convert.ToHexString gives FF0080.

How do I use Windows-1252 or another code page on .NET Core?

Modern .NET ships only UTF-8, UTF-16, UTF-32, ASCII and Latin-1 by default, so Encoding.GetEncoding(1252) throws NotSupportedException: No data is available for encoding 1252. Register the code pages provider once at startup with Encoding.RegisterProvider(CodePagesEncodingProvider.Instance), then Encoding.GetEncoding(1252).GetBytes("€é") gives 80E9. Windows-1252 and Latin-1 differ in the 0x80 to 0x9F range, where 1252 has € and curly quotes.

How many bytes will a string take, and can I encode without allocating an array?

Encoding.UTF8.GetByteCount(text) returns the exact size, 10 for "naïve ☕", and GetMaxByteCount(text.Length) returns a safe upper bound, 24 for those 7 chars. To skip the array, encode into a span: Encoding.UTF8.GetBytes(text, buffer) writes into a stackalloc or rented buffer and returns the count. Encoding.UTF8.TryGetBytes (.NET 8+) returns false instead of throwing when the buffer is too small.

What is the difference between Encoding.UTF8 and new UTF8Encoding()?

Encoding.UTF8 writes a BOM when used by a StreamWriter or File.WriteAllText (the file starts EFBBBF) and replaces invalid bytes with U+FFFD. new UTF8Encoding() and new UTF8Encoding(false) never write a BOM, which is also what File.WriteAllText does when you pass no encoding. new UTF8Encoding(false, true) also throws DecoderFallbackException on invalid bytes. For GetBytes itself there is no difference: neither adds a BOM to the array.

Run it yourself

Open any of these in the full C# editor to tweak, run and share.

C# playground