Wednesday, May 25, 2011

Curating a Discovery Environment

[Update: 11/8/11: Book published last week].  Late last year, my good friend David Swords of EBL asked me to contribute the opening chapter to a forthcoming book entitled "Patron-Driven Acquisitions: History and Best Practices." This collection of essays and research studies will be published in July by DeGruyter, and includes contributions from a range of librarians, vendors, and publishers on this hottest of topics.

Samuel Johnson aside ("none but a blockhead ever wrote except for money"), the discipline of writing always leads to learning, and with luck a good concept or two. As I thought about the changes in collection development and management that have taken place in the past decade, it struck me that the work of selectors now emphasizes:
  1. Collecting for the Moment: given the growing availability of book content in electronic form--and the growth of secure digital archives such as Hathi Trust, it is no longer necessary for individual libraries to collect for the ages. Rather, the task is to assure that content can be delivered at the moment it is wanted--from wherever it may be. This links "collections" (or more properly, access) much more closely to discovery.   
  2. Curating a Discovery Environment: this is now the central task of the activity formerly known as collection development. Or, as outlined in the chapter I have called "Collecting for the Moment: Patron-Driven Acquisition as a Disruptive Technology":
The philosophical shift underlying PDA is profound and multi-dimensional. Instead of curating collections of tangible materials, libraries have begun to adopt a new role: curating a discovery environment for both digital and tangible materials. Instead of deliberately trying to identify titles most relevant to curriculum and research interests within a discipline, broad categories of material that may be relevant are enhanced for optimum discoverability, immediate delivery, and partial or temporary use. Instead of purchasing materials just in case a scholar may one day need them, PDA offers “just in time” access to needed titles or portions of titles. Instead of collecting for the ages, libraries are using PDA to enable more targeted collecting for the moment.
This same logic applies to deselection. As we begin to come to terms with digital books and widely-shared print collections, it is critically important to assure that items removed from the library shelves remain discoverable and deliverable through other means. We continue to curate the discovery environment, and to enhance delivery options. Paradoxically, it may actually make sense to enrich the metadata associated with deselected items. This might include not only descriptive metadata, but also "availability" metadata: a URL to a Hathi Trust public domain copy; print holdings in a shared archive; commercial e-book availability; print-on-demand; or links to used book dealers.

Monday, May 16, 2011

The Cost of Deselection (10): Summing Up

The Project:  To remove 10,000 volumes from a 250,000-volume library to make room in the stacks for 1-2 years' collection growth.

Assumptions: Librarian time @ median rate of $35/hour (including benefits) and 40-hour work week. Batch data comparison for 250,000 titles would cost a minimum of $10,000 to prepare, trouble-shoot and execute. (This is essentially an educated guess, based on logic outlined here and here.

The Process and Estimated Costs:

     Project Design and Management:            100 hours        $ 3,500
     Data Extract                                              20 hours          $    700
     Develop Deselection Criteria                     100 hours         $ 3,500
     Communication w/Stakeholders             100 hours         $ 3,500

     Data Comparison w/ 3 targets (batch)                              $10,000
     Title Review from List only                                                  $ 5,810
     Disposition & Record Maintenance                                     $ 5,000

===========================================================================
   
     TOTAL with list-only review                                $32,010
     Price per volume:                                                     $3.20

     TOTAL with in-stack review                                $39,000
     Price per volume:                                                     $3.90

     TOTAL with staged review                                   $41,400
     Price per volume:                                                     $4.14


The numbers are easy to poke holes in, so please feel free to do so. We have laid them out as fully as we currently understand them, but they will vary with local conditions. There is still a great deal to learn here. We are interested in other perspectives and experiences. In most cases, the per-volume cost will decrease substantially with volume--i..e., deselection criteria, once developed, can be applied to many more volumes for only incremental cost.

The steps in the process may also vary somewhat, but there are really none that can be skipped entirely. Unless a library manager is willing to forgo communication with stakeholders, to prevent review of candidate titles, or to act based solely on usage data, the deselection process will require meetings and discussions, decisions and adjustments. These take time, and probably more time than estimated here.

These are intended as reasonable estimates, and as a starting point for discussion. Comments are welcome and revisions are likely.

Links to related posts:

Monday, May 9, 2011

The Cost of Deselection (9): Data Comparisons Revisited

Time to exercise our perpetual beta clause. In trolling through this growing string of posts on deselection costs--with an eye toward toting them up--it struck me that the entry on "Data Comparisons" could benefit from an attempt at more specific estimates, both for manual searching and batch matching of candidate files to comparator data sets. In the example outlined in that post, the task was to compare 300,000 candidates to three targets: WorldCat, CHOICE, and Resources for College Libraries--essentially, to perform 900,000 searches, then record and compile the results.

This is an area where costs are very difficult to estimate. Some libraries may be able to forgo the task of assembling deselection metadata that extends beyond circulation and in-house use, but probably not too many. Withdrawing monographs can be controversial. Evidence of holdings in other libraries, print archives or in digital format is essential to making and defending responsible deselection decisions. Comparing low-circulation titles to WorldCat holdings is the most helpful first step, especially since the holdings data gleaned there can include Hathi Trust titles, some print archives, and specified consortial partners. At minimum, all low-circulation titles in the target set should be searched in WorldCat.

There are a couple of ways to do this. Manual searches based on OCLC control number, LCCN, or title offer one approach, but it is time-consuming to key the searches, interpret and record the results. To date, the smallest data set SCS has worked with included 10,000 candidates; the largest 750,000. At 60 searches per hour (1 minute each), it would require 5,000 hours to search 300,000 titles. At $8/hour for student workers, the cost would be $40,000; at $20.68 for a library technician, the cost would be $103,400. Searching 750,000 items in this manner, besides equating to cruel and unusual employment, would require 12,500 hours. Clearly, this is not a realistic option--so much so that we won't even calculate the time or money involved.

The obvious answer, for a candidate file of any significant size, is to use computing power for what it's best at -- automated batch matching. Over the past six months, my partners and I at Sustainable Collection Services have worked closely with library data sets ranging in size from 10,000 titles to 750,000 titles. We have learned a great deal about how to prepare data for large-scale batch matching, and how to shape results usefully. Among other things, we have learned that there are many variables in both source and target data sets.

We have worked with clean data and not-so-clean data. We have found many mutations of control numbers. We have learned to identify malformed OCLC numbers, and a good deal about how to correct them using LCCN, ISBN, and string-similarity matching. Our business partner and Chief Architect Eric Redman even invoked the fearsome-sounding Levensthein Distance Algorithm, which sounds cool and turns out to be not so fearsome after all. But a few snippets from our internal correspondence may give provide some local color related to data issues.
1. Matching by title: I have implemented a basic starting point that can be enhanced as we gain experience.  This is by using the very simple Levensthein Distance algorithm. We can plug in more sophisticated algorithms later. With this basic starting point I am able to determine which LCCNs are suspect and which OCLC numbers are suspect. When only one is suspect on a record, I can use the other to correct the suspect value.
2. Using this approach, I can detect [in this record set] a pattern of a leading zero of the OCLC number having been replaced by a "7." This occurs in 1621 records. I can correct this.
3. I should begin to catalog these patterns of errors with LCCNs, OCLC Numbers, and ISBNs. I can create some generalized routines that we can plug in depending on the "profile" of the library's bibliographic data.
Or, as we began to extrapolate our response to these anomalies, and to build a more replicable process, which we would embed into SCS ingest and normalization routines:
1. Do a title similarity check on OCLC Numbers by comparing the title retrieved from WorldCat via OCLC Number to the title from the library's bib file.
2. If step 1 yields a similarity measure of .3 or less, where 0 is a perfect match, then stop. Otherwise, use the LCCN from the bib file to do a WorldCat title similarity check.
3. If step 2 yields a similarity measure of .3 or less then stop. Otherwise use the OCLC Number from the step 2 WorldCat rec to replace the library's OCLC Number.
4. If step 3 did not yield a match, log the library's bib control number and OCLC Number for manual review. The manual review is where we'll identify patterns like the spurious 7 in the Library's initial file.
This particular exchange actually concerned the smallest file we've yet handled, a mere 10,000 records! As we fix these errors in the data extract provided to us by the library, it seems clear that an opportunity also exists to provide some sort of remediation service to the library, improving its overall data integrity. But that's a story for some other day. 

Our learning curve has been steep and steady. We have now worked extensively with the WorldCat API and learned some lessons about capacity the hard way. We have worked with target data sets that require match points other than OCLC number. We have made mistakes and we have fixed mistakes. We have made more mistakes and fixed those too. And ultimately we have managed to produce some useful and convenient analyses and withdrawal candidate lists for our library partners. We have performed multiple iterations of criteria and lists in order to focus withdrawals by subject or location. In short, we have done what any individual library would need to do in order to run similar data comparisons.

None of these is a trivial task, and everything gets more interesting as files grow larger. SCS has found that effective data comparison requires top-flight technical skills for data normalization and remediation. It requires configuring and launching multiple virtual processors to improve matching capacity. It requires creation of a massive cloud-based data warehouse, already populated with tens of millions of lines. It requires high-level ability with SQL and a frighteningly deep facility with Excel. Formatting withdrawal lists for printing as picklists is also much more time-consuming than one might expect. Because SCS is doing this work repeatedly, we continue to learn and improve. We have developed some efficiencies and will continue to do so. But most individual libraries will not have that advantage. If you plan to do this in your own library, also plan on your own steep and steady learning curve.

Which, finally, brings us back to the questions of costs. Automated batch matching is clearly the way to go, but automation and creation of batches may require significant investment and a surprisingly wide range of high-end technical skills. These are not cheap. While it remains difficult to know what dollar figure to ascribe to data normalization, comparison, trouble-shooting and related processes, it is clearly substantial.

Links to related posts:


    Tuesday, April 26, 2011

    The Cost of Deselection (8): Disposition Options

    The past seven posts have brought us to the point where a decision to deselect has actually been made. Depending on whether the library has opted for in-stacks review or staged review, the books to be withdrawn or stored are visible in the stacks or identified in the staging area. As some participants might frame it, "the intellectual work has been done." What remains is to follow through on the decision and make the final disposition of the items--and maintain the records.

    Even here, however, there are several possibilities. Among the libraries with which SCS has worked, we have heard the following possibilities discussed. Withdrawn books can be:
    • transferred to storage
    • donated to another library
    • sold
    • shredded
    • recycled
    Each option carries its own particular requirements. And of course each also carries its own costs. In all cases, though, the best first step, especially when dealing with a batch as large as the 10,000 in our example, is to suppress the bibliographic records from display in the catalog. This creates the opportunity to manage the batch of records and volumes as needed. Suppression can be done as a single batch process once a withdrawal list has been converted to a picklist, with barcode numbers included. The cost, if handled in this manner, is small enough that we won't even count it separately!

    However, the cost for other steps can vary considerably, both for record maintenance and for materials handling. In general, costs increase to the degree that the 10,000 unit batch must be broken down into smaller portions to handle. To put it bluntly, it is much less labor-intensive to throw books in a dumpster than it is to pack them 20-25 to a box, or to fill wheeled bins for shredding. It is much less labor-intensive to dispose of an entire batch at once than to sell the portion that may be saleable over time. Let's look at each option individually:

    • Transfer to Storage: For record maintenance, transfers lend themselves nicely to batch work, in which changes to location, circulation status, and a few other data elements can be handled for many items in a single step. However, for shared storage, it is sometimes necessary to replace an individual library barcode with a consortial version, and to apply new ownership stamps. This necessitates handling every item individually.
    • Donating books to another library or charitable organization such as Better World Books or Folio Fund is often seen as preferable to outright discard. The books will have a second chance to circulate elsewhere. Withdrawals can be managed in batch, as can removal of holdings from OCLC. But the books have to be boxed for shipment, which requires packing approximately 500 boxes plus the expense of shipping them to their destination. Not all books are eligible for donation, which in some cases requires searching to determine if titles qualify. 
    • Selling withdrawn titles also has a certain appeal, and if we think too hard about how much money we actually spent to acquire these 10,000 books in the first place, it's tempting to try and recoup some of it. In our experience, this is dangerous ground. In her 2005 study "Library book sales: A cost-benefit analysis", Audrey Fenner determined that all of the prevailing approaches cost more in staff time than the revenue they generated. There are some new tools on the market that could significantly improve prospects here, such as Alibris/Monsoon Commerce Solutions, which help assess the value of batches of titles and offer options to support selling. (These will be addressed in another post later this spring.) But absent new tools, selling should be approached with great caution.

    • Shredding is in most cases unnecessary, but some libraries elect to go in this direction--at least for smaller batches--to avoid the visibility of a dumpster full of bound books. In the examples we have seen, titles to be shredded are put into smaller container that hold 200-250 volumes. While this is less labor-intensive than boxing them up, it constrains batches somewhat and of course the bins have to be transported to the shredder. Record maintenance options are essentially the same as for transfer, sale, or donation.

    • Recycling, especially at scale, minimizes materials handling. Cartloads of books are simply emptied into a large container once the record maintenance has been completed. The transaction costs are low, but such a vessel certainly increases the chance of complaints from faculty, library staff, and visits from the student newspaper. Without proper advance communication, these uninformed responses from the community can end up absorbing more time than has been saved.
    At this point, we are finally getting back to the email threads that kicked off my original post about deselection costs. University of Arizona Libraries estimated that post-decision steps for deselection cost them approximately $1.00/volume. UCSD's number was approximately $.25-$.50/volume. Those estimates included record maintenance and reliance on student workers and batch processes. In UCSD's case, the books are apparently sent to Surplus Sales, relieving the library of responsibility for conducting the sale, and (perhaps?) being transported without having to be boxed. In U of A's case, it is not clear what disposition options are invoked. This may be an area where it's harder to generalize about costs, or, as suggested here, different disposition options may be being pursued. But it seems reasonable to assume a minimum cost of $.50/volume for the most straightforward of these processes, and significantly higher for processes involving boxing titles up--for whatever reason. We are nearing the end of this trail. In the next post, I will try to pull together all the strands into a complete model ... and target for twittering calumny.

    Links to related posts:

    Wednesday, April 13, 2011

    The Cost of Deselection (7): Staged Review

    Although the comic possibilities are rich, "staged" review, as used here, does not imply a tableaux arranged for an audience. (Cf. the staged meeting, such as Big Heads, or deselection as spectator sport). Alas, the term is used here in the materials-handling sense, as in "moved to a temporary location."  Before all else, this requires finding a temporary location, itself not the easiest of tasks on most campuses. Locating the necessary "swing space" within the conveniently-located main library is even less likely.

    1. Finding Swing Space: To accommodate the 10,000 withdrawal candidates in our example will require substantial space, especially if they are to be handled all at once. The standard library cantilever-style steel shelf unit is 36" wide, and contains either 6 or 7 shelves, depending on whether the top shelf is used.

    At full capacity, each shelf holds approximately 30 books. Each bay, then, holds 180-210 books. In order to avoid required hard-hat use, let's stick with 6 shelves and 180 books.
    • 10,000 books will require 56 shelving bays for staging. That equates to 168 linear feet, or 84 feet if units are placed back-to-back in the common shelving configuration. That's more space than most libraries will have available, except possibly in a storage facility.
    • If instead we break the 10,000 books into five more digestible chunks of 2,000 each, the space needs become more manageable: 12 shelving bays, 36 linear feet, or a double-facing row of 18 feet. To give some idea of scale, the shelves pictured to the right will accommodate roughly 1,440 books, if all four bays both back and front are used. 
    • But while smaller sequential batches alleviates the space problem, it also creates five separate review cycles instead of one. Each cycle must be managed, and unless two such spaces can be found, each review cycle must be completed before the next can be started. If the space is in a remote location, it is conceivable that an individual deselector might need to make five separate trips to that facility. This could be known as "breaking their will", but will more likely become "ticking them off."
    2. Picking Withdrawal Candidates: "Picking" books is also used in the materials-handling sense--finding the volume and removing it from the shelf--rather than "choosing" it, since that has already been done via our candidate list. Once swing space is available (we'll ignore those costs for the time being), picking (and grinning) can begin. In a recent project conducted in a well-run storage facility with the collection recently inventoried, managers estimated a picking capacity of 500 books/day, using 2 students with light supervision. This includes finding books from a picklist, dealing with missing books and other errors, loading to carts, moving the carts to the swing space, and arranging books in call number order. At 500/day, it would require 20 person-days to move all 10,000 books. Each day's labor would cost $200, at $10/hour for each student, plus some allocation for supervision -- a total of $4,000 to pick the books and stage them for review.

    3. Librarian/Faculty Review: This process is very similar to the in-stack review described in the last post, except it is more convenient for the selector--and subsequently for record maintenance and final disposition. The actual review is similar enough to use the same numbers. Selectors can review 250/day. For 10,000 books, that equates to 40 selector-days. Selector days cost $280 each at the median rate of $35/hour. Total for selector review: $11,200. The total cost for staged review:
    • $4,000 for picking
    • $11,200 for deselection decisions
    • $15,200 total
    • $1.52/volume

    Additional considerations which apply to both in-stacks and staged review:
    • Review space will need to be equipped with a workstation, wireless access, or mobile capability to facilitate look-up of some titles.
    • Review "windows" will need to be enforced, especially if the process is broken into multiple parts. Vacations and schedule conflicts might force longer review periods. 
    • A default decision should be applied to all unreviewed titles after the specified date.
    Selector review--A Rough Comparison:  Remember, the costs here include only selector review. The books are still sitting on a shelf.  We have not yet integrated the previously calculated project management and data comparison costs. We'll save that for the grand total!
    • Review from list: $.56/volume
    • Review in stacks: $1.28/volume
    • Staged review: $1.52/volume
    Links to related posts:

      Monday, April 11, 2011

      The Cost of Deselection (6): In-Stack Review

      For some reason, deselection decisions seem more difficult than selection decisions. The thought processes are similar, but the consequences are different. But are they really? In selection, the worst-case scenario is not selection of a book that is never used, but rather non-selection of a book that is later wanted but no longer available. The best-case scenario is selecting a book that is ultimately used many times before anyone knows that it is wanted. In deselection, the worst-case scenario is the removal of a book that is subsequently wanted but no longer available--very similar to the worst-case scenario for selection. The best-case scenario is removal of books that will never be used, and retention of those that are subsequently used.

      But there remains a perceived finality to the deselection decision. That perception is in many respects false, as it is perhaps easier to re-obtain content now than at any time in history. Irreversible mistakes are rare. But the stakes seem high, and for now it is rare indeed to meet a selector who will withdraw a book without looking at it. In the previous post, we considered deselection from lists. Here we look at the costs associated with book-in-hand review. There are two approaches, each with slightly different time investments: in-stack review and staged review. Today we consider the first of these:

      In-stack review:

      Some rights reserved by Andrew|W
      In this model, books are reviewed in situ. While it is possible for a selector to use a list arranged in call number order for this process, it is more common that staff or student workers find and mark withdrawal candidates in the stacks, either by applying colored tape or turning the books down. This carries the risk of users disrupting some indicators, but also confines searching and error correction to hourly workers.

      1. Marking the Candidates: Again, let's assume that a student worker, using a list arranged in shelflist order, can find an label 1 title per minute. This may seem slow, but this average has to incorporate looking for items that are missing or misplaced, annotating the list, and moving or marking found items. In eight hours, 480 books could be marked. Let's round up to 500/day.

      To mark our entire list of 10,000 withdrawal candidates, then, would require 20 person-days at 500/day. If we assume an 8-hour day and apply a student wage of $10/hour, the process of finding and marking the withdrawal candidates is about $1,600.


      2. Librarian/Faculty review: Once the candidates are marked, librarians or faculty members can enter the stacks, either with or without a list, to review and decide on withdrawal. While finding them will be easy, the decision-making will almost certainly be slower than the marking process. Deselectors will look at context, at nearby items that cover related topics. They may choose to consult the catalog to see what other books are held by the author. They will need to indicate their decision in some way, and perhaps provide a reason for retention.

      Let's assume these thought processes and actions take twice as long as the marking; i.e., allocate an average of about 2 minutes per title. That equates to 250/day. Our 10,000 titles would require a total of 40 person-days. If we assume an 8-hour day at $35/hour, the cost of in-stack review by deselectors is approximately $11,200. ($280/person-day x 40 days).  

      The total cost for in-stack review:
      • $1,600 for marking
      • $11,200 for deselection decisions
      • $12,800 Total, or
      • $1.28 per title
       Links to related posts:

        Monday, April 4, 2011

        The Cost of Deselection (5): Title Review from Lists

        The previous steps in the deselection process have been directed at identifying candidates for withdrawal. In our hypothetical model, we have focused on titles that:
        • have not circulated since 1998
        • were published in 1999 or earlier
        • show more than 100 US holdings in WorldCat
        • did not appear in either CHOICE or Resources for College Libraries
        These are fairly conservative parameters, but as noted in previous posts, a final deselection decision creates an irrepressible urge to double-check. This is completely understandable, but it is not completely free.

        In some deselection scenarios, these candidate titles might be destined for storage rather than withdrawal. This tends to make the decision less fraught, and there may be less need for review. (It may also create the need to revisit the decision years later, but that's another story.) But let's be strong, and assume that we have a list of 10,000 withdrawal candidates that are actually intended for withdrawal. Most libraries, at least for now, feel the need to review these titles, and call on subject librarians, the Head of Collection Development, or teaching faculty to perform this task.

        Title-level review of candidate lists is quite possibly the single most expensive step of the deselection process. How expensive depends on how the work is designed and executed. The key variables are:
        • Will the review be conducted from lists or from book-in-hand?
        • Will book-in-hand review be performed in the stacks, or will books be separately staged?
        • If duplicate copies are involved, will both copies be inspected to find the one in best condition?
        Today's post addresses the first option, in which qualified deselection candidates are reviewed from a list. An example of such a list is pictured here.


         In this case, the list includes live links to both the library's OPAC and WorldCat, along with location, call number, and other details. Remember that these titles have not circulated for more than 12 years, and that more than 100 other US libraries hold them. Also remember that every hour spent reviewing such a list by a librarian costs $35/hour. Teaching faculty hours are presumably even more expensive, but do have the virtue of not drawing from the library's budget!

        Let's project a review rate of 1 minute per title. This average allows time to scan titles and groupings, plan an investigation strategy, and to review some portion of the candidates in context, by drilling into the OPAC and WorldCat. The process is in some respects similar to review of approval plan materials, but in reverse. To review 10,000 titles at 1/minute requires 10,000 minutes, or 166 hours. 166 hours @ $35/hour = $5,810. For one person, this would represent just over 4 weeks' work. To look at it another way, this list-based process costs just over $.58/title.

        In practice, the 1/minute rate may be optimistic, and the proportional costs actually somewhat higher. And of course there is the question of opportunity cost as well: what else could those 4 weeks have been used for? Still, list-based review, particularly when the candidate list is based on criteria agreed in advance, clearly costs less than physical review. We'll consider those costs in the next post.

        Links to related posts: