Quotes from Slaporte ()
“ I haven't figured out how to add text to the djvu's hidden layer, so the files do not have a text layer yet. Is Hathi's text preferable to just running it through OCR (e.g. any2djvu or tesseract) ? Is it worthwhile to try to sync the text of EOs you already have on WS to the pages in these compilations? I will look for a guide on adding a text layer when I have a chance to tinker with it, and I can upload the Hathi text layer somewhere if that is helpful for you. Also, it is relatively simple to remove pages from the djvu, so I can take care of that when we get the other big issue solved. ”
“ The text of the EOs we have for 1936 to 1979 are few and far in between at best. What we have access to (listed a little higher up from the Hathi table) is plain text from SGML files at the same archive as the USSC project gets its case opinions from (1948 on up to 1980 something) . That method is pointless, not to mentioned doomed for deletion from WS, now that we can get page scans of the original, hopefully with some sort of text layer to work from instead, Plus both scans and text from 1929 to 1948 which are missing from the public domain for the most part anyway. ”
