“ It seems that people are going to do whatever they think best for each project. Some will maintain linebreaks, while some will fix spelling — most will do something in between. Is that a fair thing to say? I prefer using the unicode characters in the text, and not templates, just because it's so much easier to write (and read whilst editing) . But I take back what I said before about trusting the Unicode Consortium to determine what characters we can and cannot use — I am currently transcribing Terræ-filius, and I want the st (that's a long s joined to the top of the t) ligature! ”
Unicode
Definition and stakes
Unicode is a universal character encoding standard that enables consistent representation of text across all major writing systems, replacing fragmented legacy encodings with a unified framework. Authors like Marie Lebert highlight its role in ensuring platform-independent character uniqueness, while John Woldemar Cowan emphasizes its ability to support global scripts.
Jim Tinsley notes its efficient transformation formats (UTF), and Thomas Stanley underscores the importance of UTF-8 in modern software. By aligning with ISO/IEC 10646 and extending beyond it, Unicode ensures interoperability, influencing digital communication and software development worldwide.
Quotes about “Unicode”
Marie Lebert, The Internet and Languages [around the year 2000…
“ First published in January 1991, Unicode "provides a unique number for every character, no matter what the platform, no matter what the program, no matter what the language" (excerpt from the website) . This double-byte platform-independent encoding provides a basis for the processing, storage and interchange of text data in any language, and any modern software and information technology protocols. Unicode is maintained by the Unicode Consortium, and is a component of the W3C (World Wide Web Consortium) specifications. ”
John Woldemar Cowan, The Complete Lojban Language — Chapter 17 - As Easy As A-B-C? The Lojban Letteral System And Its… (1997)
“ Computerized character codes Since the first application of computers to non-numerical information, character sets have existed, mapping numbers (called “character codes”) into selected lerfu, digits, and punctuation marks (collectively called “characters”) . Historically, these character sets have only covered the English alphabet and a few selected punctuation marks. International efforts have now created Unicode, a unified character set that can represent essentially all the characters in essentially all the world’s writing systems. ”
