Warning: Constant ABSPATH already defined in /home/wordsend/public_html/wp-config.php on line 29

Warning: Cannot modify header information - headers already sent by (output started at /home/wordsend/public_html/wp-config.php:29) in /home/wordsend/public_html/wp-content/plugins/wordpress-mobile-pack/inc/class-wmp-cookie.php on line 50
digital [humanities|libraries] – Page 6 – Words' End

Readex: Aguera and Sweeney on National Digital Newspaper Program

Helen Agüera is in Program Development with the US National Endowment for the Humanities (NEH). Mark Sweeney is in Preservation Planning at the US Library of Congress (LC). The full title of their joint presentation is: “National Digital Newspaper Program: Enhancing Access to America’s Newspapers.”

Agüera first. NEH programs: preservation & access; scholarly research; education; and public programs.

US Newspaper Program (USNP): started in 1980s; grants to do inventory & catalog all newspaper holdings within a state, and do selection microfilming. This is a partnership between NEH and LC. Its accomplishments: over 140K newspaper titles, 70 million pages of newsprint in microfilm. Every state is included (to varying degrees?). NEH funding totaling over $54 million was necessary to complete this program.

New program to provide enhanced access to newspapers by digitizing certain titles already preserved in microfilm. There’s no single library that has a complete newspaper collection, so this has to be a distributed effort in order to create a geographically representative collection. This is again a partnership between NEH and LC. LC will develop and maintain American Chronicle to make digitized papers freely accessible. This is a We the People project.

NDNP features: 1836-1922 (public domain only). Complements other dig. resources for earlier historical period. Begins chron. coverage with early 20th century and expans to earlier decades to achieve broad geogr. representation. Repurposes USNP bibliographic information for users to locate newspapers in analog formats (microfilm and print). (Interesting! So this is not an effort to eliminate paper. That’s great.)

Development phase began in May 2005: six projects digitizing a min. of 100K pages (each!) published in CA, FL, KY, NY, UT and VA from 1900 to 1910. LC contributes titles from its collection, aggregates all information, and creates a preservation framework. Prototype launched in September 2006 and the test bed results are [being?] evaluated by all partners.

No one knows what the optimal way to preserve this digital data will be. The optimal way to deal with that is to proceed in phases, and evaluate often. (Hooray for project management. We need more of that in the acad. humanities.)

Future directions: they’re planning to make awards to state projects with partners that have access to negative microfilm and digital infrastructure. (Collaboration is encouraged!) Successful projects will have an advisory board assisting in selecting titles. Titles should reflect political, economic, cultural history of the state, and have a significant chronological span (some continuity). Special consideration given to “orphan” (unavailable in digital form, papers no longer published, no recognized owner) titles. NEH awards will cover the costs of selection, digitization, and delivery of information to LC.

Mark Sweeney now, on preservation planning and [first] user interface.

Preservation is crucial for access. Their guiding principles:

– aggregate, serve, and preserve; do so consistently with missions and philosophies of NEH and LC (open/perpetual access to public; preservation of the assets that NDNP builds;

– demonstrate good use of taxpayer money);

– phased develpment (develop incrementally, keep door open for new options)

What’s open mean to them: freely accessible; available to use and re-use; persistent identification to support citation; open technical formats; interoperability and modular architecture; open-source software. (YAY.)

The only thing that is certain is change. Change in technologies available; change in user expectations; change in preservation models.

How do we plan for the future? Content is more important than today’s system. Design “system” to be expandable and interoperable with other systems; explicitly incorporate a dev’t phase.

Practical concerns: out-of-the-box solutions have preservation challenges, so they’re doing a lot of from-scratch development with an eye to making it easy for future generations to modify it. They’re building on LC’s expertise and experience today with metadata formats. They expect to learn from their awardees.

They’re aiming to interact with different archival and data needs, and have tools to ingest, manage and distribute the data.

They distinguish between information object and data object. Info. object: original newspaper or microfilm. Data object is the digital surrogate (interesting use of the word -vz): TIFF, JP2, PDF, OCR‘d text, structural metadata etc. Their archival master format is TIFF; their production master format is JPEG 2000. PDF is the derivative (end-user-oriented?) format.

More information about NDNP (including a lot of information on its technical specs) can be found on its website.

The prototype interface beta is available in the LC newspaper reading library. It’s behind a firewall, but from what I understand they expect to release it into the wild by end of January or beginning of February 2007.

Q&A. Mark mentioned that their OCR (which they show you upon request! that’s cool) is uncorrected/unproofread. I wonder if they use it in searching, and if so, how they account for inaccuracies? Mark says: currently it’s not a requirement that the participants correct the OCR, although some correct headlines and important stuff like that. The assumption is that significant words are going to appear multiple times; so if you bomb on recognizing the first couple of occurrences, there’d still be a good chance of it being recognized. They’ve thought of asking the reader community to help with proofreading, but that’s not part of the current development phase. That’s fair enough: the enterprise is huge, this is a beginning phase, and they’re already doing a marvelous job.

Mary Molinaro from Univ. of KY: with OCR, they’re finding the need to strike the balance between good and good-enough. It comes down to the quality of the microfilm, but the OCR technology is actually really good, and is good enough. From their perspective, it wouldn’t be worth it to go back and correct it all. Newspapers are very challenging digitization subjects.

Mark: newspapers are great because they’re so interesting to so many people. This newspaper repository is the first true digital repository that LC has built. On several different levels, this program is a prototype for other such programs, with different digitized objects.

Other questions are about the interface prototype, so I won’t reproduce them since I can’t show you the prototype itself. :) If you’re at one of the institutions developing the initial phase of this program, you’ll be able to see it around mid-October.

Readex: Imholtz – on the irreality of the past and its consequences in the digital world

This is the fourth annual Digital Institute. It’s remarkably relaxed. Sorry for lack of attributions in some places below: I can’t see everyone’s nameplate from where I’m sitting.

August A. Imholtz, Jr. is the Vice President of the Readex Documents Division, and an engaging speaker. This (along with all of the other talks) is moderated by Meg Meiman of American University. I should mention that this is primarily a librarians’ meeting; I feel quite out of my element and at the same time insanely intrigued.

In her opening remarks, Meg is quoting David Seaman (sic?): “we need the mutant book.” The digital book is (should be) more than a paper book, as regards functionality.

Relic == something you create and have no expectation that someone wil look at it. Record == intended to be used in the years to come. This is what Seaman Andy Mink understands primary and secondary sources to be, respectively.

Imholtz up now. He’ll be talking about the Polish philosopher Leszek Kolakowski‘s “What the Past is For.” First, two short quotations (reproduced as accurately as I was able to given my typing speed):

3rd century BC: “It is the same thing that can be thought and can be. So, what is not and cannot be, cannot be thought. The past cannot and cannot be, and therefore it’s a delusion that we can have such a concept as the past.”

1950: Elizabeth Enscone (sic?): “If the truth of statements about the past depends on present criteria, then, as present criteria may change, the truth of the statements about the past can change. But as the past cannot change… the truth of the statements about the past must lie in the past.”

In 2003, Kolakowski delivered a lecture at the Library of Congress, titled “What the Past is For.” A transcript of this acceptance speech for the first Kluge Prize for Lifetime Achievement in the Humanities can be found here.

We open with a reading of large chunks of this speech, which I’ll leave you to read yourselves (it’s not that long). Here’s the last couple of paragraphs, though, which I’ve a feeling will inform the rest of the talk:

The upshot of my remarks is modest and banal: although the legacy of myth is certainly an important and fertile source in human culture, we must defend and support traditional research methods, elaborated over centuries, to establish the factual course of history and separate it from fantasies, however nourishing those fantasies might be. The doctrine that there are no facts, only interpretations, should be rejected as obscurantist. And we must preserve our traditional belief that the history of mankind, the history of things that really happened, woven of innumerable unique accidents, is the history of each of us, human subjects; whereas the belief in historical laws is a figment of the imagination. Historical knowledge is crucial to each of us: to schoolchildren and students, to young and old. We must absorb history as our own, with all its horrors and monstrosities, as well as its beauty and splendor, its cruelties and persecutions as well as all the magnificent works of the human mind and hand; we must do this if we are to know our proper place in the universe, to know who we are and how we should act.

One might ask what is the point of repeating these banalities. The answer is that it is important to keep on repeating them, again and again, because these are banalities we often find it convenient to forget; and if we forget them, and they fall into oblivion, we will be condemning our culture, that is to say ourselves, to ultimate and irrevocable ruin.

AI would argue that the past and history are different things, and that there is no history. Herodotus, the first person to use the word in the West, said that history is really research (?), a summary. There is no past, only reconstructions of it. There is no history, only our interpretations of it. There are as many histories as interpreters of history. We create huge digital databanks to recreate history itself. (I’m not sure how much I agree with that. There’s a lot of interpretation that goes into the building of digital resources.)

Remmel Nunn’s wife is writing a children’s book on the concept of time. (Whoa. Brave woman.) One of the most interesting things in K’s thesis, according to RN, is that he’s saying you cannot find laws inside the study of history like you can, say, in thermodynamics. The point of what companies like Readex do is to take snapshots of history.

AI: The closest we can come to finding the truth value of historical statements (“Caesar crossed the Rubicon”) still amounts to belief, or faith. The point isn’t that he crossed the river; the point is, what does it mean that he crossed the river? We choose those among the various records of history (instantiations of memory) that we trust most.

Is K. justified in his assertion that we have to avoid the danger of calling everything an interpretation? How do we respond to him?

AI thinks he doesn’t prove his case. AI thinks interpretations are the only things that exist. This scares people because it moves towards undermining our ideas about Western morality.

Mary (lastname?): Agreed, there are only interpretations. We have to look at historical narratives and ask ourselves what isn’t there.

Suping Lu: “We may not get the absolute historical facts, but we always try to go to sources as primary as possible – contemporary, for example. Synthesis and analysis of those will get us to the relative truth.” Example of the beginning of the Sino-Japanese war in 1937: he actually discarded both Japanese and Chinese sources, but found the American and British (and to some degree German) first-hand accounts closer to the truth. (Americans and British at that point weren’t involved in a war, and so were neutral observers.)

How do you counter, for example, Holocaust deniers’ argument? AI says the way to do so is nonnarrative texts: images?, concentration camp cells, levels of chemical residue – things measurable in a way that doesn’t depend on narratology. This is still interpretation, but only in the sense that all language is interpretation.

AI: primary sources are relative. We delude ourselves in giving ultimate authority to primary sources (for ex., the US Constitution). Another is the case of the 1934 Ukraine famine. Question: was it a famine, or did Stalin do it? Photographs of trucks bearing grain: were they going to Ukraine, or leaving it? First-person accounts neutralized each other, and didn’t solve the question. An analysis of petrol records, however, found that the trucks were going in, not out.

Concern (Patricia ??): library collections are becoming more and more homogenized – we’re all here partly because we’re trying to capture uniqueness.

Digital resources are certainly not the Whole Picture either, they’re also snapshots. Frex: American newspapers archive does not contain all of the American newspapers, and it’s hard to know what has been lost.

In Readex bibliographic records, they indicate the political affiliation of every document’s author, because the documents have a political dimension and interpretation starts from the knowledge of where the author was coming from.

We got onto the subject of metadata, and I just asked a question to be more replied to than answered, over the next two days: what are the subjectivities in your metadata? With this I hope to bring to the fore – again – the largely unstated assumptions inherent in metadata. It’s no more objective than library categorization systems (such as the Library of Congress classification, for example).

[can’t see name] The promise of digitization is much larger exposure of whatever it is that we’re digitizing. There’s a promise of more evidence than ever before available to more people than ever before, but also a threat of homogenization of that evidence. (vz: one way to counteract that threat is to have diff., maybe competing?, sets of metadata.)

Suping Lu: our main task is preservation. We can get closer to historical truths through preservation. Amen to that.

Break time! My track record of documenting events is that this documenting drops off precipitously after the first session; but I’ll try to do otherwise. This is one amazing buncha people.

Leave it to yet another public gathering…

…to get me blogging again. I hope.

Much has happened on the personal front. For one, I’m back to the PhD gig, writing my dissertation this year. Been kinda re-evaluating this whole blogging thing, but for now I’ll just try to blog the 2006 Readex Digital Institute in scenic (whooboy, is it scenic! seriously) Chester, Vermont.

hyperlinked society, session 2

Linking in Web 2.0. Moderator: Saul Hansell, a reporter for The New York Times.

(I confess that it’s not clear to me yet what Web 2.0 is, exactly.)

Hansell: Think of the internet in terms of “cultural physics,” as a cyclotron that separates The Internet into Very Small Particles, each of which is “a piece of communication.” Based on their trustworthiness, they combine (link) into various compounds.

Nicholas Carr, former editor of the Harvard Business Review, book author, and blogger. Interested in the economic structure of “what we call Web 2.0”, in partic. how it influences how we consume media and other creative content. On consumption side, the link and other characteristics of W2 is disambiguate the units of consumption. Unit = not a newspaper or magazine but an article, for ex. On the production side, this means that each unit has to stand on its own economically (commercially) speaking.

Concern: even if you want the market to determine the above, what the hell is this market, that says everything has equal value, and that value is zero?

Martin Nisenholtz, Sr. VP, Digital Operations, The New York Times Company. Talks about writers who drive the most audience: people who write about less economically-connected topics tend to make less money!, regardless of how interesting/relevant their pieces are. Much else that I, sadly, missed.

Jimmy Wales, founder of Wikipedia talks about the uneven representation of information about certain parts of the world, as compared to other parts. The internet is actually equalizing that a bit: reporters who write about, say, Africa have way more articles than they’re able to get print-published, but they can put them on the net. Wikipedia finds the internet to be amazingly subversive, and blogs are a large part of that (see also, Iran). Ethiopia, Wales says, has gone from being an extremely journalistically open country to a monstrous Big Brother. So people have gone on the net, which has connected three groups: Ethiopians, the Ethiopian diaspora, and the larger [interested] community.

Nisenholtz again: we can throw out all communally-created artifacts and nobody would miss them. The important stuff is created by individuals. Uproar on the IRC channel! Someone asks (online) whether we should throw out the Talmud.

Ethan Zuckerman, Fellow, Berkman Center, Harvard University, disagrees with Nisenholtz more eloquently than I do above.

Discussion ensues, and there’s too much going on on IRC – I’ll try to synthesize some of the most interesting thought from the channel later on.

the hyperlinked society

I’m sitting here at the Hyperlinked Society conference in Philadelphia, blogging it alongside mindlace. I’ll probly post something from each of the six panels and conclusion that has something interesting in it. Apologies if writing is incoherent. :) Also, these next few posts are notes; reflections later.

This thing is getting audio- and video-recorded; I’ll post a link if (when?) they post it online.

First session, “Mainstream Linking.” Jay Rosen, moderator, has asked the panelists how links work in their world and what they mean.

Tony Gentile, VP of Healthline, a vertical search engine focused on health information. They participate in search engine marketing – they specifically go out to buy links from Google, Yahoo etc. So which links to buy, how to present them to the user, and what will the user see if they click on the ad link? Generally with linking there’s a feeling of reciprocity, he says, but one of the companies they had a contract with made back-linking a mandate. The link has transcended hypertext: they’ve developed an API that emulates a link structure. Also, they work to circulate people to the main areas of their site. In addition, esp. with health information, they have to be discerning as to whom they link to – and their whitelist (algorithmically and somewhat manually generated) contains about 170,000 companies.

Tom Hespos, President of Underscore Marketing, LLC, “a marketing guy” says the moderator. Hespos works with clients to get them “more linked in” – but he’s also a journalist and a blogger, and that gives him a perspective that other marketing people may not have. Points to the debut of Google as the turning point in link importance. Google brought relevance back to search; gave links an intrinsic value that they’d never had before. This might have done a bit of evil, in giving that value to links, despite Google’s motto (“do no evil”). Linking is a vote of confidence: page A linking to page B puts in a vote for page B. This is true in blogging world; businesses want links so that they can be found, vote-of-confidence has nothing to do with their attitude re: links, and that – Hespos says – should change, and quickly.

Eric Picard, Manager, Ad Product Planning, Microsoft. (Gasp! Not the Evil Empire!!1! Ohh, I’ll get over it.) Works on team called “MS Digital Advertising Solutions.” Focused on long-range planning and emergent media. He spends time trying to understand the economic model of hyperlinking – connecting people to information and people to businesses that might be relevant to that information. He started out as a multimedia designer and moved to things like VR, and then the web. He takes a broad view to the issues of hyperlinking. Thinks about the ways in which people “move through information.” Thinks about video game advertising, digital TV, areas into which people step for the first time in a commercial setting. Question is, how do we do this commercial setting (advertising) that is beneficial both to the consumer and to the advertiser – or at least doesn’t infuriate the consumer? MS, he says, should be thought of as an “ecosystem company” (?!); defends the “good job” MS is doing, supporting the “ecosystems” they work with (operating systems, for example, the MS search engine…)

Jay Rosen, “a student of multimedia,” reflects on above:

– Raymond Williams (sociologist) says in Culture and Society: “There are no masses; there are only ways of seing people as masses.” He meant that you can’t go into a northern England home and find a Mass Person. People are complicated. They don’t obey formulas. What does exist are ways of addressing people as masses. Today, all the past ways of seeing people as masses are coming apart, they no longer work so well. Now we have to specialize, and learn how to see people as a public, a community, knowledge producers in addition to being consumers. We’re good at connecting people UP – to companies, to central powers. Broadcasting is a good example of that. Today, a lot of the transformation and disruption in the media world is because the internet is good at connecting people laterally, not just vertically. The cost of like-minded people to find/meet each other has gone way, way down. If they’ve found each other, in many ways they don’t need the mass media. This radically changes the balance of power in the media world.

This became even more interesting when Rosen discovered blogs through a student who showed him Instapundit, just one link from which can instantly give an obscure blog ten thousand readers. Whoa. Rosen’s blog is PressThink, and it blew his mind that he could now write about media without having to run his writing through that same media. Holy freedom, Batman.

Q&A session I’ll leave for Ethan to describe in more detail.

NYC – small gather?

Next Wednesday the 17th I’ll be in New York, to attend the presentation of the Richard Lyman Award to one of my favorite people and an honored mentor, Willard McCarty. If you are interested in humanities computing, work in it, and/or know Willard, come show your support and appreciation, not to mention meet interesting people:

  • 3-5pm Lyman Award recipients lead a discussion of the American Council of Learned Societies’ Report on Cyberinfrastructure in the Humanities and Social Sciences
  • 5-6pm Presentation of the Award to Willard McCarty, Reader in Humanities Computing, Centre for Computing in the Humanities, King’s College London
  • 6-7pm Reception in the Library’s Margaret Liebman Berger Forum

If you’d like to attend, please call Martha at 919.549.0661 ext. 156 to get put on the guest list. (It’s free to attend, but they’d like to know how many people are coming, and if you tell them you’re working on something related, they might be able to introduce you to like-minded souls.)

I’m planning to get there… well, early-ish, not too early. Lunch, anyone?.. 1:30ish, maybe 1ish? Anywhere in the city that (a) has good food and (b) is cheap. After the event I’ll probably be tied up.

Wikiversity: vote TODAY!

Auspicious, that I only discovered Wikiversity today: voting that will determine whether the project will go ahead will end at midnight UTC!

The purpose of the Wikiversity project, which will ultimately reside at www.wikiversity.org, is to build an electronic institution of learning that will be used to test the limits of the wiki model both for developing electronic learning resources as well as for teaching and for conducting research and publishing results (within a policy framework developed by the community).

More information is at the link above. The idea needs work and much development and goodwill, but is promising. I’d be excited to participate, that’s for sure.

Please take a moment to create* a (free) new account and vote if you’re even remotely interested in this; they need a two-thirds majority to launch the beta. At the time of this writing it’s 197 Yes to 83 No, which is encouraging but awfully close.

*The interface for creating a new account is a bit misleading. Just fill in the username and password (and email, if you want) fields, and click on “Create new account.”

New book on digital history

Digital History: A Guide to Gathering, Preserving, and Presenting the Past on the Web, by Daniel Cohen and Roy Rosenzweig, is out. The online version is free; there is a print version as well, and links to online stores selling it are on the site referenced above.

I’ve only read a couple of chapters, but already it’s one of my favorite recent books. The language is measured and easy to follow. The site itself looks great, and the color scheme is well chosen. I’m no expert in historical web resources, but this looks like a great collection of fundamentals, most of which are bound to apply to other fields in the humanities. If you want a slightly disciplinary perspective on humanities computing, read this.

Internettings.

Well, I was going to write up a thorough account of the rest of the summit, but am a bit blogged out – if you’re interested, see what I wrote about it on the VHL blog. A large chunk of that is my previous post to this blog, so it’s not as long as it looks; on the other hand, I also link from the VHL post to the summit notes on TADA which are long but are also worth a read.

After the summit was over, I spent the weekend in rural VA not far from Charlottesville with a friend named Free who fully lives up to his name. It was lovely; there were three dogs, a horse and many children and friends, open-air fires at night and marshmallows and hot dogs roasted on long sticks. And I got to participate in prepping the foundation for a greenhouse. The air was quiet and clear, and the stars looked alien: I couldn’t make out any familiar constellations because most of the sky was taken up by tall tall trees, it was like peering up through a window to the multiverse.

Meanwhile, some intensely cool stuff is being documented on the net. For example, here’s a Quicktime movie of the Northern Southern lights from space and an .mpg movie of solar activity, both courtesy of NASA.

Also, omigod, a new Beowulf is coming out in 2007 or so, and Neil Gaiman has writing credits. !

Having missed the opening one-two-punch weekend of Serenity and Mirrormask, I’m very much looking forward to seeing both.

Summit, evening the first.

I’m in Charlottesville, Virginia at the summit on digital tools for the humanities. It’s almost 10:30pm, I’ve got to catch a 7:30am bus to campus for breakfast and a day of intense work, and I’m jazzed to the point of spinning in the elevator.

The flights were fun, and even flying out of Providence at the ungodly hour of 6am was rewarded by one of those perfectly cloudless full-rainbow-spectrum sunrises. I read and read materials for the summit, and am happy to report that I completed the reading before getting to the hotel. The hotel gave me a warm chocolate-chip cookie upon check-in.

Today was a reception and dinner, followed by the usual greetings and a keynote by Brian Cantwell Smith, a computer scientist from U of Toronto. The man knows how to talk. He gesticlated wildly, looking at times as if he was about to leap towards us from behind the podium. He talked quickly and intensely, and yet managed to keep most of the audience.

Here are some points he brought up, rephrased by me; any brilliance is his, flat language would be mine. I haven’t entirely digested all of this, so present it mostly without commentary. Your thoughts, though, would be much appreciated.

Digital is a trendy word, he said, second only to like on college campuses.

Descartes was a smart guy. He separated the work, or process, of understanding the world from the world thereby understood.

Controversial revelation of this talk: computers don’t actually exist. There do exist many devices, but what are they?

Around the turn of the 20th century, we discovered that we could fuse meaning and mechanism. An example of this would be us. This idea eventually gave birth to computers.

Computers aren’t anything special, and computer scientists aren’t studying anything special. Or maybe anything in particular. This is liberating: instead of a restricted domain, they have a sort of monopoly on the universe.

We (computing humanists) shouldn’t be party to propagating the dualism between the ostensible “abstract” and the concrete. A server going down loses not the representation of mail, but actual mail.

Descartes said that we should have clear and distinct ideas. But this isn’t the way the world actually works.

Maybe the tools we build are digital at the level of the bits, but what matters about them is humanistic.

Computers are a historical moment (a long one, which started in the mid-1800s and is still going) in which we are getting past Descartes.

Matter is both a noun and a verb. Material comes from matter.

Computing is allowing us to get past the temporary, 300-year divorse between matter-noun and matter-verb.

Our commitment to what it means to be human shouldn’t be ideological (“if it’s human, it’s good”).

People can be special as in worthy of study and careful consideration, not special as in this is where inquiry stops because there’s nothing more to say.

That’s it for now. If you’re interested in one or more elements of this, comment and we’ll talk (in comments). I’m too tired to attempt an actual argument. Tomorrow is another day, with more photography and some serious hyperte… I mean, blogging.

css.php