Tuesday, 27 March 2012

Exploring the impact of article archiving

Julia Wallace, plenary 4 (repository reality)

The PEER project is a collaboration between all stakeholder groups in scientific education, investigating the impact of large-scale systematic article archiving. Its participating publishers include the large publishers, university presses, society publishers - between them contributing 241 journals, including top, middle and lower tier journals across four broad subject areas. Participating repositories are similarly broad. Articles were either deposited by publishers or self-archived by authors; publishers provided metadata for all their articles, whether or not they deposited the full text. Authors were invited to self-deposit via a project-specific interface. Over 53,000 manuscripts were submitted by publishers. 11,000 authors were invited to self-archive; 170 did so.

Challenges on the publishing side:
  • Publisher workflows - extracting manuscripts at an unusual point in the workflow required changes
  • File formats and metadata schemas varied and required normalisation
  • Journals contained many different types of content
  • Some metadata didn't exist at early stages of the workflow eg DOI (some publishers updated metadata after initial deposit)
  • Some repositories wanted additional metadata beyond core elements
Challenges on the repository side:
  • Varying metadata requirements and ingestion processes
  • Struggled with embargo management
  • Author authentication
  • Log file provision
Ongoing research:
Independent from the executive members, to avoid bias - managed by a Research Oversight Group. Author questionnaires, usage analysis, interviews. Behavioural research looked at the behaviour of authors and users, exploring perceptions of green open access and expectations / concerns around repositories:
  • Only a minority of researchers associated repositories with self-archiving
  • Preference for final version
  • Authors don't see self-archiving as their responsibility
Costs
The research has also explored processes and costs - eg salary cost of peer review is $250 per article plus overheads. No economies of scale. Production costs of up to $470 per article. Platform set-up and maintenance costs range from $170k to $400k [interesting data!]. Challenges of competing with established publisher platforms.

Main outputs from research: preliminary indicators show 5% migration from publisher platforms to repositories. Continuing to explore accuracy of that across the board, and trends. Registration is free and open for a review meeting in Brussels at the end of May.

Defining value: putting dumb numbers to work

Grace Baynes, Nature Publishing Group leading a group discussion on use (and abuse) of analytics, with a good mixed group of librarians, publishers and one precious researcher.

What do we mean by value?
We're focussing today on the *relative* worth, utility or importance of something - numbers by themselves aren't that valuable; we need a context. We can measure downloads, time spent, social media influence, but just because we can measure something doesn't mean it is helpful or valuable to do so. We need to become more refined in how we are applying metrics.

What do we want to know the value of?
We talk primarily about journal value, but article-level metrics are increasingly important, as is the value of an individual researcher.

What indicators can we use to measure value?
Impact factor, cost, return, usage, meeting demand (qualitatively assessed). Usage breaks down in a number of ways and combines with other data eg to calculate cost per use - but what represents good value? Does it vary from field to field? How do you incorporate value judgements about the nature of the usage? Nature doing some preliminary research with Thomson Reuters here, looking at local cost per citation (ie comparing usage within institution to citations of those articles by authors within those institutions), in comparison to competing journals (Science, Cell, PLOS Biology). Picked some of the leading institutions in the US, and also looked at numbers of authors, number of citations across key journals. Grace throwing out to librarians - is this interesting? Would this data be useful in evaluating your collections?

Moved on to discuss the Eigenfactor - a Google PageRank for journals? Combining impact factor / citations and usage data in a complex algorithm.

Peer review - F1000's expert evaluations being turned into a ranking system.
what about truly social metrics e g Altmetrics (explore free demo from Digital Science, looking at tweets, blogs, reddits etc). Also ref Symplectic and SciVal as examples of visualising Analytics data.

Questions: as we move to OA, cost per use less important - what metrics will become more important? E.g. Speed to publication? BioMedCentral publish this for each article. How would it be translated into an easy to measure value?

Are all downloads equally valuable? OUP did some research into this; good articles would get double downloads (initially viewed in HTML, then downloaded as PDF if useful - so that conversion is one good indicator of actual value. Likewise, if people who *could* download full text but didn't bother, having read the abstract, that's a potential indicator of non-value).
But, this approximator of value becomes less reliable as the PDF becomes less popular as a format. Assuming that a download is more valuable if it leads to a citation - flawed - what about teaching value? Point of care use? Local citation gives a flawed picture.

The library experience
Big institutions have to use crude usage metrics to inform collection development, because more detailed work (as reported by this morning's plenary speaker Anne Murphy) is not viable at scale. But librarians know that an undergraduate download is less valuable than a postgrad download, in the sense of how important that precise article is to the reader. Citation too doesn't equal value - did they really read, understand, develop as a result of reading that article?

Ask users for reason that they're requiring ILL: what are you using this for? Fascinating insight into different ways that content is valued. Example from healthcare: practitioners delaying treatment until they can consult specific article. "We're sitting on a gold mine" - value of access to information services - showing impact. [Perhaps useful for publishers to try and capture this type of insight too - exit overlay surveys on journal websites, perhaps?]

Role of reading lists - can we "downgrade" usage where we know it has been heavily influenced by a reading list? Need to integrate reading lists and usage better.

As well as looking at usage and citations, there's a middle layer - social bookmarking sites and commentary can also indicate use / value and are much more immediate to the action of using the article than the citation.

What impact will changing authentication systems have? Will Shibboleth help us break down usage by user type? This is what Raptor project does - sifting Shib / EZProxy logs to identify users, but reliant on the service provider having maintained the identifiers and passing them back to the institution with usage data. Current bid in to JISC to combine JUSP with Raptor - agreement from the audience that this would be *HUGELY USEFUL* (hint, hint, please, JISC!)

Do people have the time to use the metrics available? One delegate recommends Tableau instead of Excel to analyse data - better dashboards.

Abuse of metrics: ref again to problem with impact factor using mean rather than median, and examples of when that has caused problems (a sudden leap to an impact factor in the thousand thanks to one popular article). Impact factor also cannot cope when bad articles are cited a lot because they're bad - not all citations are equal.

"Numbers are dumb unless you use them intelligently," says Grace. "We need to spend less time collecting the data, and more time assessing what it means."

Alternative metrics for journal quality: the usage factor

Jayne Marks is starting things off for us on another glorious day here in Glasgow. Usage factor vs impact factor - why will the extra metric be useful? What research has been done so far? What next?

Usage factor vs impact factor
With pretty much all journals online, and COUNTER well established and respected, we have a good source of reliable data to explore individual journals and their usage, as an alternative to citations (which underpin the impact factor). The impact factor, while widely respected and endorsed, is not as widely applicable across different disciplines (it's optimised for the hard sciences; in other areas eg nursing or political science, content can be well used and valued but not cited) and is US-centric. The usage factor will provide a new perspective, available and applicable to all disciplines, to any publisher prepared to provide the data. It will serve better those disciplines where usage of the content is more relevant than citations.

How did the usage factor evolve?
The usage factor project - sponsored by UKSG among others - sought to consider critical questions, including:
  • Will it be statistically meaningful?
  • Will it gain traction?
  • Will it be credible / robust?
Authors, editors, publishers were surveyed to assess whether such a metric would be of interest. Example data was analysed - 150,000 articles - to model different ways of calculating a usage factor (detailed report available from CIBER).

How is the usage factor calculated?
To avoid gaming, the usage factor is calculated using the median rather than arithmetic mean. (Comments welcome to elucidate that - my maths is a bit rusty!) a range of usage factors would be published for each journal (afraid I missed the detail on this while pondering the arithmetic - I think it would mean across a range of years). The initial calculation for a title would be based on 12 months of data within a maximum usage window of 24 months. A key question is when the clock starts ticking - when the article is submitted? When it's published online? When it goes into an issue? Does it matter if publishers decide this differently?

What might the future hold?
The team behind the usage factor suggests that ranked lists of journals by usage could be compiled eg by COUNTER to enable comparison. There are concerns about gaming but the robustness of the COUNTER stats and the use of the median should help to repel most attempts at gaming (CIBER's view is that the threat is primarily from machine rather than human "attack"); the project's leaders continue to consider gaming scenarios and welcome input from "bright academics" who can help to posit potential gaming scenarios ("we're not devious enough").

Work is still required on the infrastructure, for example, to understand how we can extract data from publishers and vendors. The project's ongoing work is being led by Jayne Marks (now at Wolters Kluwer) along with Hazel Woodward and a board of publishers, librarians and vendors. Thanks were noted to Peter Shepherd and Richard Gedye.

Question: shouldn't we be moving to article-based metrics? Marks: it could break down to that in due course.

Monday, 26 March 2012

Digging for the unknown know - If you have too much to read can a computer do it for you?

Breakout session with Eeke Smit and Maurits van de Graaf presenting highlights of the PRC study on text mining

Some highlights from the survey:

Requests from 3rd parties

77% of publishers who responded to the survey receive mining requests, but this is a low number per year – only 21% receive more than 10 requests per year

Requests from corporate customer and Abstract and Indexing companies are the most frequent.

Most who don’t receive mining requests are Open Access publishers or very small publishers.

Of those who receive requests 32%, mostly OA publishers, grant permission without restrictions. The rest consider the requests on a case by case depending on the purpose of the mining.

53% of publishers decline requests that would create products that compete or replace their original offerings.

31% of publishers ask for a fee if for commercial purposes.

Publishers mining their own data

46% publishers surveyed presently undertake content mining on own content. They are doing this for a number of reasons:

  • Improve retrieval of content
  • To generate better metadata and to add semantic tagging
  • To create new products

Of the 54% who don’t currently mine their own content, 36% plan to start mining in the next year.

Dilemma for publishers

On the plus side allowing text mining requests from 3rd parties can drive traffic to their content, but on the other down side there is potential that the outputs of text mining could function as a substitute to the original content. Finding the right balance is key.

Obstacles and solutions?

Lots of discussion from the publishers, libraries and technology vendors in the session about what the obstacles are and what the solutions might be.

Licensing and copyright issues were a big concern. Would sample licences help? STM have created a sample licence to try and prevent everyone knitting their own jumpers (to borrow a phrase from Stephen Abram's session this morning). Pharmaceutical companies are also developing their own model licences.

Lots of discussion about whether an aggregated solution might be the answer. Is there a role for discovery services such as Primo or Summon (other discovery services are available!) in providing an aggregated text mining service; they are already aggregating lot of data and have the relationships with publisher and libraries, could this be expanded on. For many, especially pharmaceutical companies, this wouldn't be comprehensive enough as the value comes from also being able to include their own locally hosted, proprietary/commercially sensitive data.

Interesting session prompting lots of food for thought!


Data-driven recommendations: supporting user search

Alison Brock, Open University, on the JISC RISE project to improve the library user's search experience

What is activity data?
Activity data is data captured about customers' activities. Big retailers use complex algorithms to analyse this data and create new revenue opportunities / increase customer retention. Meanwhile, publishers are trying to raise awareness of content and increase readership. Across the library sector, we're trying to become more business-like and making better use of our activity data is one aspect of that.

What have the OU done?
The OU has a wealth of resources from which data can be harvested: SAMS single sign on, EZProxy, SFX and EBSCO's Discovery system. They developed programs to extract data from these systems that could create recommendations; Basic data can extracted from the logs to give information about the user's subject area / course and the level at which they are studying - this enables recommendations about "other people on your course looked at". Bibliographic data can be cross-referenced (eg via CrossRef) to retrieve more metadata. Usage of resources increases the value of a recommendation, and users are given the opportunity to rate recommendations, which in turn adds further data to the database. Search terms are also added, so there's a complex web of data from which recommendations are generated - course recommendations, relationship recommendations, search recommendations. Users were given the opportunity to opt out (of having their data used to inform recommendations).

How have users responded?
User surveys have shown that the recommendations are perceived as useful by 2/3 of respondents. Focus groups have explored these perceptions further: students are particularly interested to know whether the sources of the recommendation were getting good marks! Postgrads wanted to save time - which search terms would deliver best results? Generally, there was support for recommendations but caution around provenance - users seemed to trust Google's recommendations / algorithms more than OU's, because they knew and didn't respect their peers. [this is a fascinating inversion of the commonly held view that people trust recommendations from people they know]. Google Analytics showed that course and search recommendations were used more than relationship recommendations. [intresting that OU restricted project's ability to get feedback from students - had to go through a panel run survey etc]

Lessons learnt
  • Users like recommendations in principle, but would like more detail about how they are being generated.
  • Activity data needs to be combined with other sources eg bibliographic data, student data to increase the value of recommendations,
  • The plan to release an open data set of anonymised data for other institutions to implement didn't come to fruition - lots of complexity around privacy

What next?
The data and code will be expanded for use in other projects eg OU Learning Analytics - data warehouse that enables OU to look at this data in a wider context. More info: http://www.open.ac.uk/blogs/RISE

Discussion: what do we all think about recommendation services? Sense tht the jury is still out; people don't necessarily have time to follow recommendations, and typical user behaviour is driven by search not browse.


Autosubversive practices: where we're all going wrong, and what we should do to get it right

Stephen Abram kicked us off this morning with a rallying call to librarians to shape up or move out. Following him on the podium is Martin Eve, an early career researcher from the University of Sussex. He has all of UKSG's "stakeholders" in his sights: academics, librarians, publishers are each working against their own interest in some way. We need to find common ground to build our future on; the current status quo is:
  • Younger researchers are reacting angrily against the current publishing system; they see a premium being charged for a service they don't value, while they can't access the content they need and they are suffering cuts elsewhere even as the serials budget continues to rise.
  • Yet the academic process is built on the currency of reputation, which is tied up with publishing. There is a "culture of credentialism" (encouraged by REF), and the "best" place to publish is not necessarily aligned with researchers' ideals.
  • Libraries cannot entirely predict what content is required and have to make best-fit decisions, for example, buying big deals - they too get the blame when academics can't access the wider literature.
  • Green open access is a stop-gap measure; it's unsustainable. OA needs to be gold, and the culture of credentialism needs to change. There are good examples already of the gold OA model working well, but "author pays" are scary words. Funders often don't understand it, and institutions at the bottom of the chain don't have the money.
(along the way, Martin lays down a gauntlet: UKSG's hospitality is too generous. How can we afford such nice treatment for our speakers when academic conferences can't? Thoughts on that welcome!)

What is each group doing wrong? publishers are:
  • Bankrupting libraries: academics aren't getting rich; why should publishers?
  • Pay walling: doesn't support the spread of knowledge
  • Blackmailing authors for their copyright (because authors have no choice where to publish)
  • Making libraries take the rap for being unable to afford content
  • Politicking to protect and outdated model of profit
Martin's tip for publishers: "you have to make your product as good as the alternatives that have no charge, and then you'll be valued for what you do".

Libraries: what will happen when you are no longer the custodian of paid for collections? Youll lose your power, and be replaced by IT contractors. If there's value in the library's role in centralising knowledge, you need to stop being passive.

Researchers: we drive library acquisition by where we publish; we have our heads in the sand. It's because of us that the system is as it is; we need to defeat REF, not destroy throw the good things about the publishing system out with the bad.

How can we move forward?

We can only improve the current system if we attack it together, not separately. Publishers have priced themselves out of the market, and the current alternative (OA) threatens the role of librarians. Can we solve bring these two together - the library as publisher? Only "those who are profiting without contributing to the system" would lose.

Question: are you proposing to reinvent the small scholarly society press?
It's a bit different. Martin doesn't underestimate the effort and expertise required in publishing process, particularly at scale. There is a role for publishers and librarians in the future, but it's not the same role. (is the issue here about publishing salaries? Martin posits a library press run by 30 people on £30k - are publishers pricing themselves out of this market?)

Question: "tear it down and build it up" - but what exists in the interim? How do we make the transition? Institution-specific publishing systems would need to exist everywhere at once or people would have less access, not more.
Martin: scale it down, rather than tearing it down. Start small: get five institutions to demonstrate the model and its potential. Once viability is proved, it can escalate.

Martin closed by quoting a Banksy image: "laugh now, but one day, we'll be in charge." An amusing but chilling way to remind all of us in the room who we're serving, and the importance of meeting their evolving needs - our system is broken, and we need to embrace, rather than prevent, change - or we will be ousted.

Open for business

Tony Kidd has ably dispensed with the growing number of pleasantries that kick off UKSG: we've thanked our sponsors, awarded our prizes, paid tribute to our founder, and been welcomed by the Lord Provost. The wifi is under strain from our most connected audience yet, and the tweet stream at #uksglive is buzzing. Now Stephen Abram has taken the stage and Todd Carpenter beside me is busy taking notes for blogging later. We're off! Keep it right here at UKSGLive ... Don't touch that dial :-D

Wednesday, 21 March 2012

On your marks, get set .. Glasgow!

I'm excited to be announcing our first team of UKSGLive reporters, who'll be bringing you reports and opinions from UKSG's 35th Annual Conference and Exhibition. Convening in Glasgow on Monday will be:
You can now subscribe to this blog via RSS or email so get your feed set up now! And please give us lots of encouragement and support during the conference so we don't feel we're blogging in a vacuum. Finally, don't forget all UKSG's other social media manifestations - something for everyone, we hope:
Do post any questions or comments below - and we'll hope to see lots of you soon, in Glasgow!

Wednesday, 7 March 2012

A new home for reports from the UKSG conference

Following the retitling of UKSG's journal (what was Serials is now Insights), the name of our old blog (LiveSerials) no longer made sense, so we're starting over as UKSGLive. Please update your RSS readers and blogrolls - and note that, finally, you can subscribe by email. We're still using Blogger, so hopefully you should be able to find your way around relatively easily.

Help wanted: be one of our bloggers
With the 2012 conference fast approaching, we are actively seeking new bloggers to help relieve some pressure on those who have done it in the past. Please get in touch with your Google account email address if you are willing to give it a try. You only need a Google account and a willingness to take / write up notes, and you don't need to do it in realtime (anytime within conference week is fine). We can help you get to grips with Blogger if you've never used it before, or even post for you if that's easiest. If you take notes anyway, please consider sharing them via the blog!

Design / template: thoughts welcome
We've incorporated UKSG's lovely new visual identity into the blog's design - hope you like it! Do let us know your thoughts (via comments, or to marketing@uksg.org) - we'll be happy to keep adding to and shifting around the blog's template for a bit, to make sure it's working for everyone.

Elsewhere in social media
The UKSG conference has a diverse presence in social media:
  • On Twitter - we tweet as @uksg, and we're recommending #uksglive as the hashtag, as #uksg is getting a bit overrun with sporty schoolchildren
  • On Slideshare - our speakers' slides will appear here during the conference, as fast as our little mouse clicks can carry us
  • On Paper.li - a round-up each day to help you keep on top of the tweeting
  • On Lanyrd - find out who else is attending
  • On Facebook and LinkedIn - do a bit of virtual networking prior to your arrival in Glasgow