Showing posts with label PEER. Show all posts
Showing posts with label PEER. Show all posts

Monday, 3 June 2013

Badly-coded affiliations: a too long-standing curse


  A webinar on the Repository Junction Broker (RJB) Project being presently carried out at EDINA National Data Centre in Edinburgh was delivered last week by Muriel Mewissen, RJ Broker Project manager. The RJ Broker is a SWORD-based tool for automated content delivery into institutional repositories which will identify target IRs by associating the co-authors' affiliations to their institution's platform (where available).

In the course of this RSP-organised event, Muriel shared some slides with an analysis of the preliminary content transfers the RJ Broker has performed so far. The first RJB-mediated transfer test involved processing in excess of 60,000 Europe PubMed Central articles and delivering them into the (mock) worldwide repository network.

EuropePMC is a solid disciplinary platform for the biosciences, whose content is often delivered straight from publishers. The platform's contents do usually feature good-quality metadata as a result, and EuropePMC provides thus a good example for testing research article transfer. Moreover, the specific EuropePMC article set selected for this test was remarkably modern. However, the statistical figures Muriel presented for the RJ Broker's ability to resolve author's affiliations in EuropePMC articles were simply astonishing (see figure below): author's affiliations were badly coded for over half the transferred articles' metadata.



This is a well-known issue the PEER project also had to deal with at the time. Institutions have been telling their authors since ages to try to harmonise their affiliation when signing their papers, but it's still very frequent to find affiliations such as Department of Psychology, Compton Rd or Radiology Unit, Hearts Lane which are literally impossible to process by the RJ Broker since they lack their main affiliation node.

A large collective effort needs to be done in order to provide the means for somehow tackling this long-standing issue once and for all, and ORCID looks a very promising initiative in this regard. If it were somehow possible to have author's affiliations coded into their ORCID iDs – something ORCID is actually aiming to do – the rate of miscoded affiliations could be expected to rapidly drop as a result.

Very much like the author identification, this is of course a huge challenge no-one has so far been able to tackle, and ORCID faces a lot of hard work in order to find a way to attack the miscoded affiliation issue. But there is currently much talk in the community about organisational IDs and having some system put in place that will hopefully provide the means to start solving this seemingly unsolvable difficulty. The research information management community badly needs ORCID to succeed in this challenge if it is to be able to ever start building the eagerly awaited service layer on top of the infrastructure one.




Thursday, 9 May 2013

Are publishers "the enemy"?


  This interesting issue came up (again) at the Author ID Tutorial delivered within the 4th COAR Annual Meeting in Istanbul - and it might be useful to devote a couple of reflections to it here. This Author ID Tutorial was jointly delivered on May 8th by Titia van der Werf from OCLC and myself as part of an attractive set of four tutorials at the COAR event - with the selected topics for the tutorials being as good a hint on the way things are evolving around repositories as the workshops themselves.

ORCID was big on the Author ID tutorial, to the extent that the timeschedule for the activity had to be updated on the spot in order to make room for the large number of questions and reflections prompted by the ORCID presentation. I'd like to address one of these questions more thoroughly here, namely the reluctant attitude some very qualified colleagues show towards ORCID due to the fact that the initiative seems very much publisher-driven - this making it probably not that interesting for the scholarly community.

This is again about the antagonism between publishers and the academia, and about whether both communities may at some point overcome such antagonism - real or perceived, it does not make much difference - in order to jointly work for pursuing a common benefit. This discussion is certainly interesting since it goes to the heart of a critical issue that has traditionally prevented a deeper implementation of Open Access, namely the fact that both publishers and Open Access community see each other as "the enemy". Mike Taylor - to mention just one inspiring example - regularly writes in an eloquent fashion about the reasons why the scholarly community may consider publishers to be the enemy of knowledge dissemination. However, same way as a certain degree of (informal) agreement was reached at the COAR event that the fight between advocates of Green and Gold OA is a pointless diversion of energy and will only harm their common objective, it could very much be argued that making emphasis on the differences and the misbehaviours over the good practices in collaboration may result in blocking win-win cooperation opportunities.

It is true that publishers such as Elsevier and databases such as the TR Web of Science or Scopus are a big driver behind ORCID - although the fact that over 130,000 researchers worldwide have chosen to individually register their ORCIDs as of May 3rd should not be overlooked either. It is evident too that a widely implemented successful persistent author identifier scheme will benefit publishers very much - but it will benefit institutions and especially authors even more. There was again an agreement at the author ID tutorial that this is something that needs to be done, and when examining the wide range of previous attempts to achieve the goal of author and work identification and disambiguation, it becomes clear that having publishers involved in the initiative provides it a significantly larger chance of succeeding.

I have repeatedly written here about the encouraging effort the EC-funded PEER project did in bringing together publishers and Open Access repositories and how advisable it would be to try to further explore opportunities for collaboration - some of which are indeed being exploited, see for instance Wiley's direct involvement in the JISC-funded PREPARDE project for research data publishing. ORCID is certainly one of these opportunities and with all due respect to constructive dissent, it would be an exercise in shortsightedness to let it slip away.




Friday, 22 March 2013

... and preaching to the non converted


  After quite a long time away from this blog due to various circumstances – with work overload probably being the most convincing one – I will try to catch up with various threads in the next day ot two – and I shall start the attempt with an answer to this request for explaining what Open Access is and what its aims are I was delivered from the interesting comments section of this “Whoops! Are Some Current Open Access Mandates Backfiring on the Intended Beneficiaries?” post by Kent Anderson at The Scholarly Kitchen blog. This is my answer – I tried to keep it as concise as possible, apologies if it may still be a bit long.


I am probably too busy trying to overcome the numerous challenges that stand in the way of Open Access implementation myself to provide a too detailed and accurate description of what Open Access is and what its aims are, but I'll give it a go. Let me start by quoting the Berlin Declaration on Open Access to Knowledge in the Sciences and Humanities (2002):

"The Internet has fundamentally changed the practical and economic realities of distributing scientific knowledge and cultural heritage. For the first time ever, the Internet now offers the chance to constitute a global and interactive representation of human knowledge, including cultural heritage and the guarantee of worldwide access".

According to this, Open Access means ensuring this possibility is realised, and worldwide dissemination of research outputs should indeed be a shared goal for institutions (and its libraries) and for publishers. It means that any researcher anywhere in the world may have the opportunity for the first time in history to freely share her research results (and this includes research data) with the whole research community and beyond. Whether this is achieved through the so-called Gold route (Open Access or hybrid journals) or via Open Access repositories (the Green route) is secondary to some extent - although not of course if business models are our sole concern here.

Open Access deals with the have and the have-nots (which does not just mean developed vs developing countries, but rather privileged vs underprivileged researchers in terms of having or not an institutional coverage for accessing the research information they require for carrying out their own research). And Open Access deals with whether a freely available author's final peer-reviewed manuscript might provide a useful alternative to the much-preferable version of record for those underprivileged researchers who can't or won't afford paying the fees required to read the papers that will allow them keep up-to-date with advances in their own research area.

Research funders are well aware of the challenge, especially those in the area of biomedical research, and Open Access mandates are their attempt to tackle the access issue in an area where many institutions both in rich and poor countries lack the (quite substantial) budgets required to provide their reseachers a comprehensive access to publications in toll-access journals. What about publishers? They are indeed adapting their business models to fit the Gold route by taking Article Processing Charges from authors as a prerequisite to making research papers available Open Access so they can meet the funders' mandates – which is fine. But this adaption to Open Access has not at all improved their image in the eyes of institutions (and many researchers in them), who suspect some not-so-subtle form of double-dipping is taking place since they still need to pay for their journal subscriptions on top of the APCs.

What could publishers then do to stop the fight?

The European PEER Project was a 3-yr STM Publisher Association-lead attempt to assess the impact of Open Access repositories on the 'European Research ecosystem'. This was technically carried out by delivering a large amount of final peer-reviewed author manuscripts into a cross-European institutional Open Access repository network.

Publisher participation ensured the right research article version was deposited, and the whole exercise was also useful for them: not only they were able to become aware of the relevance of sufficient metadata (a concept that CrossRef has later extended among the wider publisher community), but also to harmonise their interoperability standards through the use of the NLM DTD. Furthermore, the conclusions of the PEER project assessment carried out by CIBER Research Ltd was that not only publishers were not harmed by Open Access repositories, but rather on the contrary the paper download figures from journal pages at publisher websites were much improved by their availability as final manuscripts at repositories (since it's the version of record any researcher will prefer to read and cite unless of course they have no means to accessing it).

PEER was a one-time exercise, but it also delivered a proof of concept for cooperation between publishers and institutions in order to provide researchers the service they require for meeting the funders' mandates they are subject to. And in fact some sensible publishers are still interested – and taking subsequent steps in this direction – in delivering their authors the deposit service they require to meet the mandates. The way these sensible publishers see it, this is a means to offer researchers competitive advantages at journal selection time and will ensure a steady number of submissions in an increasingly competitive market framework for journals.

In the meantime the institutional Open Access community (which reached a critical mass quite a long time ago) is taking steps to ensure the repository systems become fit for purpose in order to meet funder requirements in terms of offering OA to the outputs of research projects funded by them. There are indeed technical as well as cultural/political challenges, in fact quite a number of them, but there is also a sustained and persistent effort to figure out the best ways to gradually address them. Institutional Research Committees are suddenly becoming aware (and this is the concern comment #2 addresses) that institutional research publishing budgets won't reach for providing Gold Open Access via payment of APCs for the whole institutional research output, so they're instead turning their eyes to their institutional Open Access repositories and wondering whether it could be the way of meeting funder mandates in a much cheaper fashion. At the same time, some funders are starting to rule hybrid journals out of their mandates for compliance purposes on order to avid the abovementioned risk of double-dipping.

The landscape keeps hastily evolving and it seems further adaption will be required both from publishers and institutions. This could ideally happen through cooperation and not through struggle, but there seem to be too many prejudices and too little efforts out there for a constructive dialogue to take place in a sustainable way.


Friday, 1 June 2012

PEER End of Project Conference: a few reflections



  The fact that the PEER European Project (Publishing and the Ecology of European Research) has managed to establish a fruitful communication channel between publishers and repositories was repeatedly highlighted along the PEER End of Project Conference held last Tue May 29th in Brussels. This ability for fostering a successful collaboration between stakeholders initially at conflicting positions is undoubtedly one of the main PEER outcomes and it would be good news for the Open Access movement as a whole if these communication channels could remain open in the future. As Norbert Lossau put it, favouring pragmatism over ideology could be very useful for jointly outlining evolving business models.

The second most important achievement of the PEER project was being able to establish a tested publisher-repository transfer infrastructure which can be deployed beyond the project. A good number of PEER components and technical findings -such as the PEER Depot dark archive, adoption of the TEI format as an unique metadata interchange standard or SWORD as standard transfer protocol, the way usage is dealt with or the use of the GROBID component for automatic metadata extraction- are potentially re-usable for other ongoing or future publisher-driven transfer initiatives and especially valuable for automatic item transfer into repositories within an hegemonic Gold Open Access scenario that was also frequently predicted along the meeting.

Additional publisher-driven deposit initiatives such as Japanese 'Zoological Science meets Institutional Repositories' were mentioned along the conference as well as COAR involvement in the interoperability strand pottentially offering opportunities for follow-up work. Besides that, the JISC-funded SONEX Group has repeatedly underlined along its analysis of deposit use-case scenarios the strong workflow similarities between PEER and the JISC Open Access Repository Junction (OA-RJ) Project carried out at EDINA in Edinburgh. The RJ Broker feature -which performs a very similar role to the PEER Depot 'moulinette'- is currently being enhanced and will shortly be offered as a service through the UK RepositoryNet+ Project.


The figures associated to the PEER project are certainly impressive: 53,000 stage-two manuscripts (aka post-prints in SHERPA RoMEO terminology) from 241 journals published by 12 mainstream publishers were processed by the PEER Depot resulting in 22,500 EU manuscript deposits (including embargoed papers) released into six different IRs plus into a long-term preservation archive at the KB in The Hague. Two submission routes were designed: automatic publisher-driven deposit and 11,800 invitations to authors for self-archiving their papers, the latter one resulting in just 170 author deposits (or 0.2% of total PEER deposits).


The large difference between deposit figures associated to the two deposit routes led PEER researchers to conclude that authors sympathise with OA but don't see self-archiving as their task, therefore "Green OA not being the key road to optimal scholar information systems". The PEER Usage research -one of the three research team projects within the PEER Research strand along with Behavioural and Economics research- proved also that although current findings reflect the position of a relatively early stage in PEER development, Open Access repositories are not really a threat to publishers (thus confirming the so-called "no effect" publisher hypothesis). In fact, making pre-prints visible in PEER repositories actually generates more traffic to publisher sites, although the ever growing rates of publisher downloads make it hard to supply an accurate measurement of the impact on publishers of post-print availability in repositories. Ian Rowlands from CIBER Research Ltd estimated that publisher full-text downloads increased by 11.4% as a result of earlier version of papers being available at the IR.

Gold vs Green OA

While testing Green Open Access and its economic consequences for the publishing ecosystem in Europe was the main PEER goal and Green OA was the preferred workline when the Project started back in Sep 2008, the Gold Open Access route seems nowadays to be winning hearts and minds of those trying to promote access to research output on a wide basis. PEER has produced quite a number of evidences on the fact that Green OA does not harm journals nor publishers, but in the meantime attention has shifted to Gold Open Access and hybrid journals as a way to ensure that final publisher/PDF versions of the papers are made available.

This is probably the strongest argument in favour of Gold OA, but there are also very good ones that support Green OA. As a result, a lively debate is taking place these days inside the Open Access community on which OA model should receive main support from the government bodies. Many voices argue as well that both models should co-exist, as the research output coverage will be wider as a consequence. And there is finally an important fact to be accounted for after watching PEER result of 99.8 vs 0.2% automatic vs author-driven deposit: author self-archiving rates should not be systematically used as reliable indicators of the strength of Green OA, since there is nowadays a wealth of alternative ways to populate repositories that do not imply self-archiving obligations for authors. In fact CRIS systems, their integration with IRs and the resulting alternative workflows for content ingest into repositories were not mentioned at all last Tuesday despite having already been proved effective by a recently released UKOLN report. When trying to offer a fair estimation of Green OA relevance based on the wider deposit picture, the contribution to repository population from these alternative workflows should also be considered.

Thursday, 15 March 2012

Sobre PEER y awareness-raising


  Publica Ángel Borrego (Departamento de Biblioteconomía y Documentación de la Universitat de Barcelona) en el Blok de BiD un extenso e interesante comentario sobre el informe final "PEER Behavioural Research: Authors and Users vis-à-vis Journals and Repositories". Siguen a continuación algunas consideraciones adicionales al respecto, quizá un tanto extensas para incluirlas como comentario al post:

Además de etiquetar las versiones de los trabajos especificando claramente si se trata de la versión publicada o de alguna clase de versión intermedia -algo que se está haciendo cada vez más sistemáticamente- los repositorios institucionales harían bien en distinguir con claridad su sección científica de la de 'otros materiales académicos' (incluyendo por ejemplo fondos patrimoniales). Esto es algo sobre lo que primero DRIVER y más adelante OpenAIRE han hecho notable hincapié, tratando de identificar repositorios con infraestructura científica para el Espacio Europeo de Investigación. La perspectiva desde las bibliotecas universitarias no suele ser sin embargo tan unánime al respecto, y sería quizá ahí donde habría que comenzar una labor de difusión eficaz de lo que son y pretenden los repositorios institucionales.

Por otro lado, en relación con los conocimientos sobre acceso abierto y repositorios de los autores/investigadores, así como sobre su valoración de los mismos como herramientas de difusión de su producción científica, queda claramente mucho camino aún por andar. No obstante, muchos repositorios han alcanzado ya el suficiente grado de consolidación como para poder presentarse como una sólida infraestructura científica institucional en los congresos científicos, facilitando así su conocimiento por parte de los investigadores (véase por ejemplo este 'Computer applications and quantitative methods in Archaeology 2012' que se celebrará próximamente en Southampton con un notable énfasis en aspectos relacionados con el acceso a las publicaciones, sobre todo en el ámbito de datos de investigación).


Finalmente comentar que además de estos aspectos relacionados con el awareness-rising, el proyecto PEER tiene interesantísimas cuestiones que debatir respecto a la interoperabilidad de repositorios y las oportunidades y los retos técnicos que plantea la transferencia de contenido entre plataformas (en el caso de PEER desde plataformas de editores hacia repositorios). Con COAR y el Grupo SONEX trabajando ya sobre estas cuestiones de interoperabilidad, la Conferencia fin de proyecto de PEER del próximo 29 de mayo en Bruselas puede ser una excelente oportunidad para una nueva entrada sobre PEER, sea en el propio Blok o en algún otro foro.