java - invalid byte 2 of 2-byte UTF-8 sequence

Question

Welcome To Ask or Share your Answers For Others

java - invalid byte 2 of 2-byte UTF-8 sequence

1 Answer

深蓝 · Answer 1 · 2021-10-23T18:40:17+0000

Most commonly it's due to feeding ISO-8859-x (Latin-x, like Latin-1) but parser thinking it is getting UTF-8. Certain sequences of Latin-1 characters (two consecutive characters with accents or umlauts) form something that is invalid as UTF-8, and specifically such that based on first byte, second byte has unexpected high-order bits.

This can easily occur when some process dumps out XML using Latin-1, but either forgets to output XML declaration (in which case XML parser must default to UTF-8, as per XML specs), or claims it's UTF-8 even when it isn't.

Categories

java - invalid byte 2 of 2-byte UTF-8 sequence

java - invalid byte 2 of 2-byte UTF-8 sequence

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags