RMIT University
Browse

Separating the wheat from the chaff: Identifying key elements in the NLA.au domain harvest

journal contribution
posted on 2024-11-01, 14:19 authored by Geoff Fellows, Ross Harvey, Annemaree Lloyd, Robert Pymm, Jake Wallis
In 2005 and 2006 the National Library of Australia (NLA) carried out two whole-domain web harvests which complement the selective web archiving approach taken by PANDORA. Web harvests of this size pose significant challenges to their use. Despite these challenges, such harvests present fascinating research opportunities. The NLA has provided Charles Sturt University's POA (Preservation for Ongoing Accessibility) research group with access to these web harvests and associated keyword indexes. This paper describes the 2006 harvest and uses the example of blogs to address how to identify material within the harvest and determine issues that need further investigation.

History

Journal

Australian Academic & Research Libraries

Volume

39

Issue

3

Start page

137

End page

148

Total pages

12

Publisher

Routledge

Place published

United Kingdom

Language

English

Copyright

© 2008 Taylor & Francis

Former Identifier

2006043880

Esploro creation date

2020-06-22

Fedora creation date

2015-01-19

Usage metrics

    Scholarly Works

    Keywords

    Exports

    RefWorks
    BibTeX
    Ref. manager
    Endnote
    DataCite
    NLM
    DC