He fucking did it again.
One un-smashed window remains in our car.
He’s some punk kid. I saw him but, of course, not his face.

searching for the ineffable
He fucking did it again.
One un-smashed window remains in our car.
He’s some punk kid. I saw him but, of course, not his face.
Some bored soul has busted in the rear windshield and one of the back door windows of our car tonight. Probably with one of those little hammers that come in safety kits: the windshield is broken in the middle and also around the edges, but not in-between – and there are no marks on the metal around the edges, so it probably wasn’t a sledgehammer.
Fucking punks. Completely pointless vandalism. Unless they know us and are trying to make some sort of a statement, which I highly doubt: nobody really knows us around here, except by passing hellos. I hate this sort of shit, this mindless mean-spiritedness. [ETA: They didn’t take the ipod that was in the glove compartment, or my passport, or the road atlas. Really pointless.]
Now, instead of angsting about my dissertation tomorrow, I get to call the insurance company and sweep up broken glass, and maybe drive the car over to some repair place or other (or maybe wait until they get the part in, since these cars haven’t been sold in the States for that long). Plus there’ll be a deductible-sized hole in our budget that we just didn’t need. All of it because some idiot broke stuff for ha-has.
Work on the thesis is being done in spurts. Partly of necessity: I’ve managed to schedule myself for five out-of-town events, in four trips, all within the same two months. Plus there’s a possibility of getting to spend time with the Nephew, who is turning into an excellent if willful little person, and while I like having him over, I’ve discovered that zero work gets done when he’s around. I suspect it’d be different if we lived closer together, but so it goes.
Anyway, my mental pattern with regard to thesis work has been repeating – rather predictably so; this pattern has existed since long before Roland. It goes something like this:
Before the work period starts: attitude cavalier, anxiety far at bay. Usually, during this time things are happening that make me feel good – conferences, family time, just-breaks with good books.
First day of work period: attitude of “ok, here I am buckling down.” Permitting myself to spend this one day Organizing, which never takes just a day, so the day almost inevitably ends in a vague state of many things accomplished but not enough, damnit.
Second day: overwhelmed and in denial. Repeat mantras of the “I’m smart enough and diligent enough to do this in time – but it won’t happen if I keep succumbing to the anxiety and denial, for they evilly drain energy, confidence and time” sort. Keep having to remind myself that I love this stuff (and I do, it’s the time pressure that’s a bitch to deal with).
From this second day on, if I manage to get myself to start working sometime before 10, life is good. If I don’t, I lose days to self-loathing, or thought patterns less dramatic but just as draining.
The hardest thing is not knowing how long this thesis will take, or how much work it will be. I’ve set myself a hard limit – graduate next spring – but what if I don’t get the work done? If I could see the steps clearly, it would be easier to work. As it stands, it’s hard to even know which large swathes of work will turn out to be useless for the current purposes. There’s no way I can do justice to the Roland corpus in the course of this dissertation; defining its limits in a field of material that I don’t know that well (there’s SO MUCH of it!) seems like a futile exercise.
Outlines don’t help, either. They take so much time to make, and then I have to change them a million times over. My working outline helps me organize whatever it is I want to do next, I guess. But since it isn’t representative of the final thesis structure (as I discovered rewriting chapter the first), it’s not an indicator of how far along I am.
These are some of the things that make thesis writing hard – I’ve read and heard this from many sources. The resultant anxiety is a pest, and I resent it for that. Go. Shoo!
Whew, that was grand. Just one thing about the closing panel discussion, while it’s fresh in my mind.
This year’s CaSTA was billed as “a joint computer science and humanities computing conference.” And it was! And [we saw that] it was good. Of the five keynote speakers, three were humanists and two – computer scientists. The final discussion was called “Humanities Computing Science??”.
William Arms, in his remarks during the panel, said that during the conference a word was frequently used that isn’t generally used in his usual [computer-science] circles. That word – knowledge. He, and just about everyone at the panel, said that what they primarily want from the “other side” is dialogue.
In light of that, what I’d like to see in this continuing dialogue is a bit of discussion of the word science. As it’s been used lately (in the last, what, 200 years?), it implies “HARD.” Humanities implies “soft.” That’s a major point of contention.
But given that “science” pretty much means “knowledge,” should we revisit our use of the word?
Ray Siemens is a computing humanist and Renaissance scholar working at the University of Victoria. The full title of his paper is “Knowledge management and textual cultures? Work toward the Renaissance English Knowledgebase (REKn, pron. “reckon”) and its professional reading environment (PReE).”
REKn seems to be aiming to amalgamate and integrate knowledge in its area. Its implementation is based in the study of disciplinary activity of, and professional interaction among, those in the humanities. It’s founded in concepts of knowledge representation and modeling. A short description of the project can be found here.
Knowledge representation: draws on the field of AI and seeks to produce models of human understanding that are tractable to computation. Modeling: REKn/PReE model data, intellectual processes, and beyond.
Key elements of REKn’s model:
– representation of archival materials
– analysis/critical inquiry originating in those materials
– the communication of the results of these tasks (the dissemination of primary and secondary materials)
REKn’s assumptions: all of the above are interrelated and inseparable, and electronically representable.
They’ve collected primary and secondary sources, and have built tools for working with them (the tool-building process seems to have been multi-stage: many tools built and discarded as inadequate). They’re looking to long-term partnerships with Renaissance materials providers in the future. Right now REKn has about 13,000 primary sources and over 80,000 secondary sources. About 1500 of these resources are currently available for public use, but the majority are not open-access.
So that’s REKn, the text base. What about PReE, the reading environment? It’s a rudimentary document viewer, and analysis and communication facilitator. Currently the UI is a “down-and-dirty prototype,” as they’ve been concentrating on making things work in the back end. [vz: he’s showing PReE in Windows; I wonder what it’s written in.] They’ve made several analytical tools for the encoded texts. Primarily, though, analysis will be carried out using TAPoR tools.
Communication facilitated electronically: they’re attempting to provide a system by which people can manage their professional interaction.
Short-term goals:
– integrate better with TAPoR and the Public Knowledge Project reading tools
– conduct usability studies
– consult with “contextual” stakeholders, including acad. publishers
– move prototype to a web environment
– scale up!
And that’s Ray’s talk, the last talk of this conference. Next is the panel discussion, titled “Humanities computing science?”. The panel will consist of the five keynote speakers. I’m not sure whether I’ll be taking notes on this; it’ll be video recorded, and there’s little probability that I’d do it justice. Again, I’ll update when the webcasts are up.
Eugene Lyman is an Early English scholar at Boston University. The full title of his paper is “Presenting an electronic critical edition of Piers Plowman B.” The B refers to one set in a multitude of manuscripts of this Middle English alliteraative poem, comprising only one of its versions.
EL frames his talk with the following two quotes:
“Computers compute, of course, but computers today, vfrom most users’ points of view, are not so much engines of computation as venues for representation.” –Matthew Kirschenbaum
Understanding the poetics and principles of electronic scholarly editing means understanding that the primary goal of this activity is not to dictate what can be seen but rather to open up ways of seeing.” –Martha Nell Smith
Lyman has created software that allows you to look at the existing manuscript pages from ten different manuscripts, enlarge portions of those pages, read the transcriptions, look at erasures that tell us interesting things about how people might’ve edited in medieval England. You can search within the dataset, visualize the text in various ways, view the underlying XML markup… yesterday (or was it the day before?) EL actually gave some of us an informal demo of this thing, and it is sweet.
How to make a continuity of presenting a single text that exists in multiple manuscripts?
EL created the Elwood Viewer, which looks at documentary editions of texts. Its aims:
– tight coordination of text and image
– visual cueing to guide/reinforce reader’s attention
– handy tools, all within a metaphorical arm’s reach
– ease of navigation, especially at opportune moments
– parsimonious use of screen real estate
– simple, no-cost programming environment, open to change [this is all JavaScript, I think… -vz]
The software allows you to extract data for further analysis. One such analysis called into question the notion that scribes were totally random with respect to the ways in which they encoded [marked up, oh yes, for markup exists in many forms including punctuation and embellished first letters] their texts.
This is all a cricital edition: you take a group of witnesses, compare them, note the variations, and choose the variations you think were on the author’s agenda when she was writing the text. The notion of a critical edition, which effectively “leaves behind” the actual manuscripts, is pretty controversial; nevertheless, at the moment critical editions have a lot of weight in literary studies. [vz: I’m leaning toward the camp that’s sceptical of critical editions, but only in the sense that I don’t tend to consider them definitive, while many others do.]
EL has a prototype version of the critical edition. I haven’t found any screenshots of it on the web, but here’s the Piers Plowman Electronic Archive hosted at the University of Virginia. It seems more a worksite than a ready-made resource, but is worth poking around.
Johanna Drucker is a professor of media studies, and a founding member of UVA’s Speculative Computing Lab. The full title of her keynote is “Graphic conventions: visualizing knowledge and subjectivity.”
Can we shift from information to interpretation, and shift [back] to a more humanistic point of view within digital humanities? To do that, we have to re-introduce subjectivity, and maybe substitute the mechanistic with the probabilistic (the latter being humanists’ worldview, according to JD).
Visual information conveys information in a form that makes it very hard to analyze systematically. Two ways to create stable knowledge in a notation system: one is with [natural] language, the other – mathematical notation, said Rene Toms of the Oulipo. He never did talk about visual representation, and with good reason: it’s an unstable notation mode.
Subjectivity comes in two forms: position (structural) or inflection (semantic).
The notion of information comes from a particular set of assumptions of what knowledge is. JD not interested of getting rid of this model of knowledge; but rather to propose another model of knowledge.
Visualization is compact, the problem isn’t having enough space to represent information – it’s pinning down the exact nature of the information and our assumptions about its aspects.
Visual information (VI) can lie or misinform, like natural language can.
The silliness of chunking of processes (how authors write stories: “author thinks about a topic” –> “author sketches an outline” –> “author reviews the sketch”) is apparent, but we have to do this sort of chunking when we’re working in a computational environment, which requires discrete units. Because of schematics’ rhetorical power, we eventually come to believe them.
So what about text visualization (as opposed to data visualization above)? Interest in doing things to and with texts is very active, especially within the creative-writing communities. Like TextArc, that sort of thing. They can be silly, ugly, destructive of the original, and yet they have their uses.
Edward Tufte, the exquisite engineer according to JD: information pre-exists visualization. Visualizations can be transparent enough to get us access to information. JD disagrees: visualizations are interpretive, opaque, distortive. They create informmation.
Temporal modeling at SpecLab. Basic assumption: timelines as they are conventionally defined and designed come out of the empirical/natural sciences. Assumptions there: time is unilinear; time is homogeneous (metric is stable); time is continuous (no unbroken intervals in temporality). None of these three things hold. Temporality branches in our lives, in poetry/film/etc. Time is not homogeneous – some moments fly by, others are long (the moment before the kiss and the moment after are very different, JD says). Time is not continous, either: there are breaks/ruptures, recorded in historical accounts for example.
SpecLab constructed a grammar of inflections, of visual elements they’d use to represent time, types of events and their relations to each other. There’s a lot of information JD is giving about what SpecLab has been doing; I’ll point you to the Lab’s site instead of summarizing.
In the IVANHOE game, every action takes place from within a role. Each role has a set of assumptions that go with it. They employed it as a teaching tool at UVA, with the purpose of showing that, in fact, every action stems from a set of presuppositions. [vz: that can’t be right, I’ve captured too simplistic a description. Go see the site for more.]
Subjective meteorology: JD’s current project. An art project, which JD says – duh, art is in the humanities. [vz: yay!] She charted and graphed and visually represented a bunch of weather patterns – lines of anxiety/anticipation, storms of anger – which look gorgeous on the slides but don’t seem to be on the web. These representations can be chained together and animated to playfully and visually represent one’s subjective perceptions of the world around.
Great discussion follows. I can’t pretend to capture it well enough; I’ll post an update when the keynote webcasts are up, and urge anyone interested to watch this one when it’s available.
David Hoover is at NYU, and is the Vice-President of the Association for Computers and the Humanities. The full title of his paper is “CaSTAing breadth upon the waters.” (“Cast thy bread upon the waters: for thou shalt find it after many days,” say Ecclesiastes 11 – hence his title.)
DH seeks simple methods for examining word frequencies in corpora of single authors, at different stages of their production. Do authors tend to start disliking (using less) words they used to like (use a lot) earlier in their careers? Or vice versa? How does an author’s vocabulary change over her production’s life? DH does a lot of statistics to try to find out.
He’s talking about Trollope, about whom I know nothing. Apparently, the 100 most variable words in his corpus are all proper names. They also all appear in more than one stage of his writing career (early-middle-late). Henry James, however, has some non-proper nouns.
DH’s project is very much in progress; he says it’ll be a while until he has something interesting to say about the evolution of writerly language. One interesting question is: if you have a writer who starts writing very young, does their vocabulary change quickly, early on? What about writers who write far into old age?
One interesting consistency in James is that he seems to have used fewer “precious” nouns as the years went on: words like coquette and tresses.
Richard Cunningham is at the Acadia University English department, and directs the hypermedia center there. The full title of his paper is “Developing digital navigation from The Arte of Navigation.”
Readers experience something different reading an electronic document as opposed to a paper one. [Glaringly obvious, RC admits.]
RC presents the Acadia Digital Culture Observatory. They have digitized a 1561 edition of The Arte of Navigation and want to observe how readers read and use it. The original text included a navigation instrument made of three concentric paper circles of different sizes (volvelles), which are to be overlaid one on top of another and rotated. Here, you can see it for yourself. (Hm, it doesn’t seem to work in Firefox on Mac in Blackletter mode; I suggest the use of Arial to be safe, or you can download Blackletter from their table of contents.) Check out particularly the navigation instrument Flash files in the “other moving images” section of the TOC, they’re fun to play with.
Ray Siemens of Victoria hosts a session of three papers related to the Text Analysis Portal for Research. First we have Geoffrey Rockwell, with “Text empires: text analysis in excess.” Shawn Day will talk about “The use of the recipe as a guilding metaphor for flexible and efficient self-guided computing instruction.” Finally, Stéfan Sinclair will talk “On data & views in text analysis.” All three presenters are from McMaster University in Hamilton, near Toronto.
ROCKWELL.
Information overload: 5 exabytes of information created in 2002. Exabyte = 1,000,000,000,000,000,000 Bytes. It’s a thousand petabytes, or a million terabytes. [Holy wow.] Spam is cheap, but reading has costs. How can text analysis help?
Why this explosion of information?
– growth in population and wealth: more money, more media toys
– multiple-media, from the photograph (1820s) to the iPod
– digitization of information and business practices: cheap creation, storage, reproduction, and transmission
Challenges to the system: what are the effects?
– experience of information overload
– multimedia shock
– narrowing expertise (because nobody can’t keep up with a broad discipline!)
– archive fever
What can we do?
– understand the problem (literary dimension to it; a problem of scale, a bibliographic problem)
– produce less? [shock! I can hear the internal gasps around the room!]
– file and (not) store smarter
– find smarter (not more) [ooh, I’ll quote him in my dissertation work! no, I cannot, in fact, process all the litcrit written to this day]
– learn to read differently
The latter two of the above are opportunities for text analysis.
Problem of scale to text analysis for finding and reading:
– heterogeneous formats and multimedia rich
– closed (“for perfectly reasonable reasons” -GR) information empires (Google) build on existing indexes or build their own
– new questions, research methods (data mining and visualization)
– text analysis tools developed for coherent texts (collaborate with data mining & HPC [high-performance computing] community)
TAPoR.2 model, Beyond Finding and Reading:
– gathering and aggregation function (working with existing empires like Google; create your own study library (myEmpire))
– mining function (clustering and classification; provoking questions, not finding)
– interface and visualization function (effective interactions for research)
DAY
They’re using the recipe metaphor to get people of different backgrounds to use TAPoR.
A recipe for self-guided instruction:
– ingredients
– steps
– glossary
– discussion
– further information
Ingredients:
– ingenuity
– a useful metaphor
– a versatile set of tools
– users desirous or willing to consider using said tools
Steps
– identify objective
– consider users’ needs
– develop case studies that describe how your tools can meet these needs
– apply a familiar metaphorical approach to engage and instruct
– deploy recipes through a wiki
Glossary
– recipe: a useful guiding metaphor that offers optimal flexibility…. [couldn’t get it, too fast]
Further Information:
Try the recipes out! (For example.)
Nice, familiar, easy concept. As Shawn is pointing out right now, super easy to engage a beginner user. This could be very useful, as well, when getting folks used to traditional humanities research methods to try, say, text encoding.
SINCLAIR
[Stéfan is the creator of HyperPo, the coolest text analysis tool ever so far.]
Generally, there’s a one-to-one mapping between tools and the data views of their results. SS has been thinking more in terms of this progression:
text -> tool -> data (TAML) -> style -> view
Among other things, he wanted to create a framework to use in teaching the development of text analysis tools in a modular way.
It’d also be nice to be able to chain tools together – you run a tool on a text, get the resultant data and feed it to another tool, and so on. This requires tools that can ‘talk” to each other, and output data in the same (or similar enough, or easily translateable) formats.
HyperPo 7.0 is coming soon!