- Search for string figlina: returns 31 results.
- Search for Early Middle Ages (477 AD - 830 AD): returns 543 results.
- Search for Giza: returns 0 results.
- Search for pyramid: returns 3 results.
- Search for Last Name=Griffith: returns 212 results.
- Commands history for Ian Milligan's wget tutorial
- 182 files downloaded in 20 minutes.
- Commands history for Kellen Kurschinski's wget tutorial
- 80 photos downloaded in 2 minutes and 32 seconds.
- Using the Canadiana Discovery Portal's API and a modified version of Ian Milligan's bash script to scrape Ottawa-related records from 1800-1802.
- Output: 1,072,458 words
- Commands history
- I tried scraping 1800-1900 but my connection was terminated fairly quickly, seems like the website thought I was a bot at that point.
- Using twarc with Twitter Apps to REST tweets to JSON format and convert them to CSV format with json2csv.
- Output for "twarc search hist3814o > search.json"
- List of IDs for the above tweets
- Commands history
- Converting an image to text:
- Using Tesseract OCR
- Using Tesseract R
- Saving those textfiles as pictures (OCR and R) and converting them again to text:
- Using Tesseract OCR
- Using Tesseract R