Summary

Xover/Archives/2020

“ There are lots of reasons. Ironically, the biggest reason the community tends to prefer DjVu is that MediaWiki's extraction of OCR text from PDF files is atrocious and much worse than its extraction of the same text from DjVu files. You can literally open the same PDF file in Acrobat and copy the text out and get bette results than MediaWiki's. But this may be at least partly due to the biggest issue for me: there's a definite dearth of even semi-decent tools for working with PDF files, especially in an automated way. ”
Source: Wikisource

Xover/Archives/2020

“ What's happening is that the wikimarkup that we write (and of which the contents of templates are part) is first parsed by MediaWiki and turned into HTML, and that HTML is then sent to the web browser that re-parses it and renders it to the user. When MediaWiki parses the wikitext it applies heuristic rules to account for the differences between what humans write and what HTML requires, in this case regarding paragraphs in text. Humans separate paragraphs with two newlines, but HTML requires a paragraph to surrounded by

...

tags.
”
Source: Wikisource

Xover/Archives/2020

“ In order to generate a DjVu I will need to have the files locally on my computer, of course, but where I download them from isn't all that important. If it is convenient to put them in a zip file somewhere I can download them all in one go that will be the easiest; but so long as they're available somewhere I can grab them without too much trouble. If what suits your workflow best is to upload them directly to Commons then do that. Just try to make sure you use a naming scheme that is predictable and consistent, and put them all in a category for the work so they're easy to find. ”
Source: Wikisource

Get perspective with Kwize: daily news enlightened by great literature