General pattern
We meet; we’ll take a 15-20 minute break in the middle; we’ll build stuff; we’ll talk about stuff.
In certain weeks, a designated person will take us through their engagement with a particular tutorial (most will come from the Programming Historian), as indicated below. There will be readings to support this engagement; I will discuss these individually with you once we divy up the work. Some will highlight uses of the approach, or perhaps issues with the approach, or could’ve usefully been improved by the approach… or… or… or. I will expect you to also clearly articulate connections in other research you’ve done, read, or courses you’re taking/have taken (this alone is an important habit to cultivate.) You can also bring digital history projects in the wild into the discussion.
(nb Even if it’s not your week to present, it will be a richer experience if you’ve given the tutorial a shot as well.)
The remaining time will run along the lines of a mini unconference. That is, I expect you to have a sense before class of things you want to work on/discuss/collaborate on. As thatcamp.org says:
at an unconference, the program isn’t set beforehand: it’s created on the first day with the help of all the participants rather than beforehand by a program committee. Second, at an unconference, there are no presentations — all participants in an unconference are expected to talk and work with fellow participants in every session. An unconference is to a conference what a seminar is to a lecture; going to an unconference is like being a member of an improv troupe whereas going to a conference is (mostly) like being a member of an audience.
I’ve had far too many seminars that felt like dreadful dreadful conferences. So, let’s give this a try. One thing that I think I would like you to discuss every session: how does this particular tutorial move us closer to the final project goal? What could we do with this? How can we open this thing up even more?
These sessions will be opportunities for the more techy to help the less, for the more theoretically inclined to help the more methodologically inclined. I will say this though:
doing embodies theories of knowing and how you do things reveals what you know. So know what you’re doing.
This schedule will be updated with who-will-lead-or-do-what-when after our first meeting. The schedule/load can/might be adjusted depending on enrollment.
Base Camp
At base camp, the climbers assemble, the gear is organized, and a route to the top is plotted out. What are the basics that we need? The larger goal of this class for me is to develop your digital history skills so that you can undertake (and expand your imaginary of the possibilities for) digital history projects of your own. If, following Brett, we see Dighist within DH, then dighist is where we use and interrogate the transformations across physical/digital forms to tell history.
What are the sequence of transformations for the Archaeology Impossible Project?
Consider the historical evidence of climbing and its various transformations:
flowchart TD
%% Define Styling Classes
classDef source fill:#e3f2fd,stroke:#1565c0,stroke-width:2px,color:#000
classDef physical fill:#fff3e0,stroke:#e65100,stroke-width:2px,color:#000
classDef digital fill:#f3e5f5,stroke:#6a1b9a,stroke-width:2px,color:#000
classDef data fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px,color:#000
%% --- 1. Original Sources ---
subgraph Original_Materials [Original Source Materials]
D([Diaries]):::source
PH([Photos]):::source
F([Films]):::source
V([Videos]):::source
PO([Physical Objects]):::source
BD([Born Digital Materials]):::source
OM([Other Materials]):::source
end
%% --- 2. Physical & Institutional Curation ---
subgraph Physical_Curation [Physical Curation & Publication]
HJ([Handwritten Journals]):::physical
L([Letters]):::physical
PUB([Publications]):::physical
A([Archives]):::physical
MC([Museum Collections]):::physical
PC([Sold / Private Collections]):::physical
end
%% --- 3. Digitization & Web ---
subgraph Web_Access [Digitization & Web Publishing]
DW([Digitized for the Web]):::digital
DF([Films Digitized]):::digital
YTV([YouTube or Vimeo]):::digital
CE([Clips Embedded in Websites]):::digital
PW([Published to the Web]):::digital
end
%% --- 4. Data Extraction ---
subgraph Data_Processing [Data Extraction & Linked Open Data]
PIR([Photos, Images, Digital Renderings of Text]):::data
OCR([Transcribed using OCR]):::data
TXT([Text]):::data
J([Represented in JSON]):::data
LOD([Linked Open Data]):::data
end
%% --------------------------------
%% Flow Relationships / Edges
%% --------------------------------
%% Diaries paths
D --> HJ
HJ --> PUB
D --> L
L --> A
%% Photos paths
PH --> PUB
PH --> A
%% Web Digitization paths
PUB --> DW
A --> DW
%% Film & Video paths
F --> DF
DF --> YTV
YTV --> CE
V --> A
%% Physical Object paths
PO --> MC
PO --> A
PO --> PC
%% Publishing to Web
MC --> PW
A --> PW
PC --> PW
CE --> PW
DW --> PW
%% Bridging Web Assets into Data Processing
DW --> PIR
PW --> PIR
%% OCR and Transformation Pipeline
PIR --> OCR
OCR --> TXT
%% Compilation into JSON and LOD
TXT --> J
OM --> J
BD --> J
J --> LOD
What’s missing? Every one of those transformations, every one of those blocks, captures a dense theoretical and methodological mass of ideas and practices. Consider your own engagement with dighist and whether or not (or how) you’ve been thinking about what happens to ‘history’ during these points of transformation. (The code that generates that view is here and you can visualize it, experiment by dropping it into the mermaid live editor).
Background Context on Digital History in Canada
At some point before the Fall Break, please read
Gaffield, Chad. ‘Clio and Computers in Canada and Beyond: Contested Past, Promising Present, Uncertain Future’. The Canadian Historical Review, vol. 101, no. 4, 2020, pp. 559–84. link
Kim Martin. ‘Clio, Rewired: Propositions for the Future of Digital History Pedagogy in Canada’ The Canadian Historical Review, vol. 101, no. 4, 2020, pp. 622-639. link
Sept 10. Getting Started
To install:
I want you to keep your notes for this class (your thoughts, your observations, your scratch pad as you fight with things) as plain text files. Not word. Not one note. You might have a nice system set up already, but for my pedagogical purposes: plain text files using the .md file extension. It is harder to kill plain text through neglect see Brett, part 5 of his essay. Some options:
- Sublime Text,
- Notepad++
- Obsidian
- Tangent
- I am also intrigued by orsn as it seems to be tailor-made for the kind of work we’re doing.
Bibliographic management:
- Zotero for research management (bibliographies, citations, pdf annotations, and note making)
- Tropy for research management of photographic materials (whether your own photos or other kinds of imagery)
Code sharing/collaboration:
- We will set up github accounts too.
Do not pay for anything. Nothing I ask you to do here should involve paying for an account or access. If you find yourself at any point this term being asked for a credit card, stop and talk to me. )
We will spend a bit of time setting up your own personal research management environment and talking about this in general; this isn’t so much a part of ‘digital history’ as ‘strategies to keep you sane.’
To do:
Once we get set up, let’s do -
-
Some of the digital literacy pieces from this tutorial platform. Nb, this is something I pulled together; the ‘time estimates’ it makes are in no way realistic and over state the amount of time a tutorial will take (currently just uses default values).
-
Perhaps some command line shennanigans: Terminus; Command Line Murders.
-
Simpkin, Sarah. ‘Getting Started with Markdown’. Programming Historian, Nov. 2015. programminghistorian.org, link.
-
Tenen, Dennis, and Grant Wythoff. ‘Sustainable Authorship in Plain Text Using Pandoc and Markdown’. Programming Historian, Mar. 2014. programminghistorian.org, link.
To mull:
-
Baker, James. ‘Preserving Your Research Data’. Programming Historian, Apr. 2014. programminghistorian.org, link
-
Heppler, Jason A. ‘How I Use Obsidian’. Jason Heppler Weblog, July 2024. jasonheppler.org, link.
-
Heppler, Jason A. ‘How I Use Obsidian Redux’. Jason Heppler Weblog, January 2026 link
Sept 17: Digital History in the Wild
We’ll consider some examples of what digital history looks like, and the kinds of questions it might ask (or permit the asking). We’ll look ahead on the syllabus and divvy up the various tutorials.
Before coming to class read:
- Catherine D’Ignazio and Lauren F. Klein, “What Gets Counted Counts” and “The Numbers Don’t Speak for Themselves,” Data Feminism (2020), https://data-feminism.mitpress.mit.edu/.
- FAIR, CARE
- Gupta, Neha, Andrew Martindale, Kisha Supernant, and Michael Elvidge. “The CARE Principles and the Reuse, Sharing, and Curation of Indigenous Data in Canadian Archaeology.” Advances in Archaeological Practice 11, no. 1 (2023): 76–89. link.
To do
Then, select two pieces from Current Research in Digital History. Read the piece. See if you can find the underlying data. Does the project make the data available? Is there a public facing website for the larger project? Is there an intersection with Public History? What about FAIR, CARE? What data ‘counts’, to whom, and why? Be prepared to talk about how the underlying data for these projects are presented, discussed, curated, and explored. List any tools/techniques that are new to you. What can you find out about how to use such a tool? (part of the challenge of doing digital history lies in what might be called ‘Dependency Hell’ - to use this, you need that; to do that you need this other thing… and so on).
What comes next
Take a look at your assigned Programming Historian tutorial. Pay attention to the requirements, and try to work out what assumptions the tutorial’s author has made about your previous experience and what you need to know to be successful with the tutorial. What are the known unknowns (as it were) for the method?
Sept 24: No Class
I am away this day for a conference. I invite you all to get together in a coffee shop, and explore the Reviews in DH and other resources to see if there are examples of the kind of work you’d like to be doing. Then, as a group, feel free to compose an email to me saying, ‘how did they do x,y,z, and can we learn that?’ I’m happy to throw things out and rejig so that we let your interests surface.
Camp One
As we progress up the mountain, things start to get real.
October 8. Digitizing
We’ll use some of the equipment from the XLab to digitize both physical and documentary objects.
Before coming to class read:
- Ryan Cordell, 2017. ‘Q i-jtb the Raven’: Taking Dirty OCR Seriously’ Book History 20: 188-225 (see also link)
- Sparrow, T., Bain, K., Kimber, M. and Wilson, A.S. 2024 Visualising Heritage: using 3D immersive technologies to innovate, document and communicate rich narratives for HS2, Internet Archaeology 65. link
- Scheinfeldt, T. 2025. ‘Handwriting Recognition Roundup’. Found History, 6 Dec. 2025 link
To do
Bring a small object that you might like to document in 3d, or some documents you’d like to scan. For 3d work, we might start with this piece of equipment. We might try LIDAR scanning the classroom or perhaps the Quad, we’ll see.
We’ll also try Crump, Jon. ‘Generating an Ordered Data Set from an OCR Text File’. Programming Historian, Nov. 2014. programminghistorian.org, link
If you have a google gmail account, we can try dropping an image into an LLM and asking for structured data in return.
Side Quest: Speaking of images, sometimes the most useful thing we can do is systematically keep track of both the images AND our annotations. That’s what Tropy’s for. And once we’ve gotten a collection of annotated images, sharing those notes, images, and annotations might be the most generous thing we could do.
Put your research photos online using a static website via Tropy (instructions here)
October 15: Data Metadata, Paradata
Ask yourself: has your research to date been sustainable or reproducible or replicable? What do these terms mean, for us? Why should historians care about this sort of thing?
There are two kinds of ‘research data’ that we could make available. There are our primary materials (and all of their related annotations and mark-up). There are our thoughts on the secondary materials we read to contextualize those primary materials. There are the data - the things we study themselves. There are the metadata - the information that contextualizes the sourcing of that data. And there are the paradata - the information that details the transformations we have done to the metadata and data and why.
To Read
- It’s over 10 years old now, but Caleb McDaniel’s Open Notebook History still is posing questions that historians largely avoid. Have a read of what several archaeologists argue about ‘open science’ in archaeology. Where are the points of intersection between McDaniel and the archaeologists? What parts of this spark joy? What bits fire up anxiety? What does this mean for you and your research? (On a somewhat related note, see this collection of previous CU History graduate degree work; what opportunities are lost here, gained here?)
To do
- Handy bit of code: Here’s a Google Notebook I made that uses something called ‘paddleOCR’ to identify text in an image and then OCR it link. There are many other options for OCR’ing text. Use this on some of the document scans from last week. Document the result also in terms of its data, metadata, and paradata.
- See what doing open notebook history through creating a ‘datapage’ could look like here and here’s the template for making such a page ourselves. Other options exist (including things like datasette.io). Find a historical dataset and create a datapage for it.
- Cosovschi, Agustín. ‘From Sources to Data: Designing a Database for the Humanities and Social Sciences with Nodegoat’. Programming Historian, Feb. 2024. programminghistorian.org, link.
October 22: Discoverability and theorizing Search
Sometimes, historical information that we might want to study is provided via an ‘application programming interface’ or API. Sometimes, it’s in a webpage that you need to parse in order to get it into a useful format.
Because of adversarial bad-actor scraping of websites by AI companies, many websites, institutions, and individuals have had to put counter measures into place to stop scraping. The materials here might therefore have to be changed to take this into account.
Nb, ‘scraping’ isn’t necessarily a bad thing; a browser works by loading data onto your own computer and scraping is a way to automate retrieval of elements on a webpage. But if it isn’t done politely (ie, at human-scale rather than machine-scale) it carries serious costs. See for instance Eric Kansa’s recent essay in Internet Archaeology.
To read
-
Sugimoto, Go. ‘Introduction to Populating a Website with API Data’. Programming Historian, May 2019. programminghistorian.org, link.
-
Underwood, Ted. ‘Theorizing Research Practices We Forgot to Theorize Twenty Years Ago’. Representations, vol. 127, no. 1, Aug. 2014, pp. 64–72. link
To do
Let’s take a look at canadiana.ca. Enter a query, and then notice the url and how it is constructed. Modify your query using the form, and see how things change after the ? in the url.
Then give this a try:
- Williamson, Evan Peter. ‘Fetching and Parsing Data from the Web with OpenRefine’. Programming Historian, Aug. 2017. programminghistorian.org, link.
When we’re dealing with data organized in a structured tree, one can parse the structure of the website, and use the Wget command to give it a go. See
- Milligan, Ian. ‘Automated Downloading with Wget’. Programming Historian, June 2012. programminghistorian.org, link.
and
- Kurschinski, Kellen. ‘Applied Archival Downloading with Wget’. Programming Historian, Sept. 2013. programminghistorian.org, link.
Badly-formed wget commands (or commands not correctly shut down) can lead to downloading an awful lot of data and can make you look like a bad-actor, which we do not want. AI companies have largely posioned the well for these methods as something we might use.
Fall reading week Oct 26 - Oct 30
I’m leaving this here as something you should read at some point. Nockels, J., et al. 2024. The implications of handwritten text recognition for accessing the past at scale, OCR & Handwritten text. Journal of Documentation 80.7, 148-167 link.
Camp 2
At this point in the climb, we’re fully invested in the process. We might have to go back down to Camp 1 if we find we’ve gone the wrong way. We might be preparing for the next leg.
Nov 7: Networks
I was a relatively early proponents of network analysis in archaeology. We’ll talk about what a network perspective might offer (hey, it made my entire PhD!), as well as perils and pitfalls. (Here’s a past mre using a network approach on Ontario history)
To Read
-
Ahnert, Ruth, Sebastian E. Ahnert, Catherine Nicole Coleman, and Scott B. Weingart. 2020. The Network Turn: Changing Perspectives in the Humanities. Cambridge: Cambridge University Press. link. (Our library, direct link). This is a short work, read the intro and part 1, dip into anything else that strikes your fancy.
-
For examples of network analysis in the wild, this issue of the Journal of Historical Network Research is great - see in particular Ruffini’s conclusion to the issue, which addresses the ‘so what’ and the ‘we knew this already’ and ‘what if we’re wrong’. This is important. On a similar note, see Lincoln 2015 on ‘confabulation in the humanities’ here.
To do
A handy tool for quick network visualizations: https://networknavigator.jrladd.com/. Here is a dataset of the index of correspondence for the Republic of Texas; knowing nothing else about the Republic of Texas, how might visualizing this correspondence network provoke new insights or questions?
-
Düring, Marten. ‘From Hermeneutics to Data to Networks: Data Extraction and Network Visualization of Historical Sources’. Programming Historian, Feb. 2015. programminghistorian.org, link.
-
Ladd, John R., et al. ‘Exploring and Analyzing Network Data with Python’. Programming Historian, Aug. 2017. programminghistorian.org, link.
-
Brey, Alex. ‘Temporal Network Analysis with R’. Programming Historian, Nov. 2018. programminghistorian.org, link. (You can install RStudio on your machine to run R, or you can change the runtime for Google Colab from Python to R like so.)
Nov 12: Topic Models and Text Analysis
Now that we’ve got a whole bunch of text, what might we do? I love the Data Sitters Club - they’re a group of scholars using a wide variety of DH approaches to understand an important book series from the ’80s & ’90s. Read about their misadventures with topic modeling here. Let’s also play with Voyant.
To do
-
Mähr, Moritz. ‘Working with Batches of PDF Files’. Programming Historian, Jan. 2020. programminghistorian.org, link. After OCR’ing pdfs, it does some topic modeling.
-
If your documents are kep as text files, give this a try instead. The topic modeling tool uses MALLET under the hood (and you can learn more about how that works and why, here.) Here’s a zip file with the text of historical plaques from Toronto that you can try fitting a topic model to. What patterns in ‘public memory’ do you find?
Camp 3
Now we’re still cycling backwards and forwards, trying to find the route to the summit, but the way is more or less clear.
Nov 19: examining images at scale
What can we see if look at vast amounts of historical imagery at once? I’ve just completed a project looking at social media and the trade in human remains (people buy and sell human remains online). A major tool we used were various neural network models trained to discriminate different classes of materials (including retraining such models for our own purposes). This week, I’ll talk about that for a bit, and we’ll think about under what conditions such approaches would be useful in your own research, and what dangers may lurk.
To Read
-
Chapter 1 in Arnold & Tilton’s book - OA version
-
Wevers, M. J. H. F., Vriend, N., & De Bruin, A. (2022). What to do with 2.000.000 Historical Press Photos? The Challenges and Opportunities of Applying a Scene Detection Algorithm to a Digitised Press Photo Collection. TMG – Journal for Media History, 25(1) link.
-
Melvin Wevers, Thomas Smits, The visual digital turn: Using neural networks to study historical images, Digital Scholarship in the Humanities, Volume 35, Issue 1, April 2020, Pages 194–207, link.
To do
- Play with Teachable Machines
- Strien, Daniel van, et al. ‘Computer Vision for the Humanities: An Introduction to Deep Learning for Image Classification (Part 1)’. Programming Historian, Aug. 2022. programminghistorian.org, link.
- Strien, Daniel van, et al. ‘Computer Vision for the Humanities: An Introduction to Deep Learning for Image Classification (Part 2)’. Programming Historian, Aug. 2022. programminghistorian.org, link.
Nov 26: Knowledge Graphs, vectors, embeddings
Language models work by expressing patterns in the training corpus as vectors in a multi-dimensional space. Why language models seem so able to do so many things is largely a function of size and speed. In the readings below I give you two pieces from over a decade ago by Ben Schmidt (a digital humanities scholar who now is at a company called Nomic) that introduced to many dh people the idea of vectors and embeddings and the things they could do. I used his code on materials from the human remains trade to understand how sellers ‘constructed’ the idea of human remains as being something you could sell. Anyway, the technology has progressed and there are uses here for historians in mapping our materials within the huge multidimensional space the corpus as a whole describes.
To Read
- Schmidt, Ben. 2015. Word Embeddings and Word Embeddings: Rejecting the Gender Binary.
- Graham, S., Yates, D., El-Roby, A., Brousseau, C., Ellens, J. and McDermott, C. (2023) ‘Relationship prediction in a knowledge graph embedding model of the illicit antiquities trade’, Advances in Archaeological Practice, 11(2), pp. 126–138. link
- Graham, Shawn. Once Upon A Time: The Behaviour Space(s) of Stories. link
To Explore
Eric Kansa’s visualization of the the archaeological materials from Poggio Civitate, where the descriptions are expressed in a vector space and then visualized using a forced graph algorithm (direct link to viz; takes a bit of time to load).
To do
- Build a knowledge graph embedding model: colab notebook (save a copy to your own gdrive and then work from that, remember.)
- Build a custom image search model (go to google colab, open a notebook from github, and paste in this url) https://github.com/shawngraham/pn_notebooks/blob/main/2_experiment_2_Use_ArchaeoCLIP_in_a_notebook.ipynb
The Summit!
At the summit, we have achived the main goals, but there’s still some work to do. For one thing, you’ve got to let people know that you’ve made it and why it matters…
Dec 3:
Being able to control your own space online enables a certain kind of freedom. Take a look at some academics’ scholarly websites- Kathleen Fitzpatrick; Jason Heppler; Chantal Brousseau; Tim Sherratt; Jim Clifford; Kim Martin. What unifies them? How are they different? What audience(s) do they serve? What constitutes effective presence?
To read
Please read the following tutorials about building static websites, especially the why of it all. I’m not a fan of Jekyll - I find it frustrating to use - but I want you to know these things. Don’t worry about trying to put together a Jekyll powered site using these tutorials (unless you really want to).
- Visconti, Amanda. ‘Building a Static Website with Jekyll and GitHub Pages’. Programming Historian, Apr. 2016. programminghistorian.org, link.
- Visconti, Amanda, et al. ‘Running a Collaborative Research Website and Blog with Jekyll and GitHub’. Programming Historian, Nov. 2020. programminghistorian.org, link.
- Marwick, Ben, et al. ‘Packaging Data Analytical Work Reproducibly Using R (and Friends)’. The American Statistician, vol. 72, no. 1, Jan. 2018, pp. 80–88. link. What bits are applicable to us?
To do
We’ll build a website using Pelican, which is a python package that will read a folder of text files (in the markdown format), pass them through a template, and spit out the necessary html files that make a website. We’ll then put these files online using Github Pages. (See my notes here).
Or we can try cb-essay which builds a story-driven site from a google sheet filled with metadata.
You’re also welcome to give my work-in-progress ‘Polybius’ a try; like cb-essay, it’s meant to be a generator for a data-driven story telling website. Critiques welcome. See also my rendering of the the Historic Places dataset here.
And there’s a whole lot more that could be done; check out Epoiesen.
Dec 11: Sunsetting a digital project
To read
- Arts & Humanities Research Computing. Sunsetting link.
- Holmes, Martin, and Joey Takeda. ‘From Tamagotchis to Pet Rocks: On Learning to Love Simplicity through the Endings Principles’. Digital Humanities Quarterly, vol. 017, no. 1, May 2023.
- Endings Project, Principles: link
To mull Perhaps digital projects should be allowed to die? And: just because something is digital in nature, does that automatically mean that it has to be accessible to the web? Especially in this age where culture is being taken to pieces, and monetized through llms?
To do
Let’s talk about your own research/experiment. What were the data? What were the transformations? What emerged from this engagement? How you can bring your digital work to a close when doing your thesis or MRE.
Into the future?
Have a look at “The Historian’s Desktop” or “Autarch”. Taking those as indicative models of future engagement with computational approaches to the past, which of those would you prefer, as a scholar? Consider Doctorow on Reverse Centaurs.
Things are changing fast. If a single prompt can accomplish (seemingly) so much, why have I spent so much time trying to teach you the ‘hard’ way? The last thing I want you to do - and this isn’t graded, and it’s not for me to pass judgement on as an instructor - is for you to write down your own manifesto for dealing with AI, the digital, going forward.
Set out your own guidelines. Revise them, revisit them, as you go forward. Whatever you do next: be intentional. Be informed. Be clear about what you will do, and what you will not.