Wednesday, 31 August 2011

Progress Update for August

The team continue to make good progress.

Data Sources
All the new data sources originally identified have been added to the HALOGEN database.

Tools Development and Evaluation

Development of prototypes of the ‘data extraction tool’ (using Business Objects) and a ‘web based enquiry tool’ continues.

The second iteration of the ‘web based enquiry’ development was presented to Nottingham users in late June and to the Roots of the British/Diaspora researchers at Leicester in July. Feedback from these sessions has been used to enhance its presentation and functionality.

A prototype of the business objects based data extraction tool has been delivered to researchers at Leicester and they are evaluating its use.

Other

A visit from David Flanders - JISC Programme Manager occurred on the 2nd August and highlighted a number of hot topics (most notably data licensing issues and ideas on how to test our deliverables). The feedback from this session was very positive.

Photo below - left to right: Olly Butters, Andrew Bradley, Jonathan Tedds and Dave Carter.


Monday, 8 August 2011

MySQL server and ArcGIS

Olly and Liam got to grips with linking ArcGIS and MySQL server last week, essentially they have created a method to allow ArcGIS to talk directly to the HALOGEN data without the need to export data (e.g. as a csv or tab delimited text file). So far it looks like we can query directly on the database. Why is this great news you ask? Before I created an events theme in ArcGIS which was then converted to a shapefile, unfortunately with large volumes of data we exceeded the maximum size of shapefiles and therefore could not query the data any further without a crash. We still need to test and see if relational databases work though, watch this space...

Andrew.

Tuesday, 2 August 2011

Plotting aggregated points over Google Earth satellite images

Olly has created a web interface that plots our data on to Google Earth satellite images. Our data is aggregated to the centre of BNG 1km squares to preserve confidentiality and to standardise the resolution of our database as there are several different data sources to compare. I am concerned that end users may forget / not read the project documentation and think that a point marks the exact location of data when in reality it could be anywhere in a km square around the point.  Does anyone else share these concerns - or have a way of reminding an end user of this?

Andrew.

Tuesday, 5 July 2011

Progress Update for June

Good progress is being made and we have now delivered the new data sources as planned!

Data Sources

A key target was to complete the load of the new data sources to HALOGEN during June and this has been achieved. Well done Olly and Andrew !

The coverage of the PAS data has been increased to all of England. The Capelli data and 1881 Surname census data has now been added to the database.  

The load of an additional source of surname data has been requested and will be added over the next few months.

Tools Development and Evaluation

Development of prototypes of the ‘data extraction tool’ (using Business Objects) and a ‘web based enquiry tool’ continues.

A plan for the iterative development of the ‘web based’ enquiry has been agreed. The second iteration of development is now complete and the system was presented to users at Nottingham University's Institute of Name Studies on 28th June.


Other

A communications plan to support internal dissemination activity relating to the project has been drafted for Board approval.

We had our first Project Board meeting on 27th May and that went well.

Dave Flanders, our JISC Programme Manager has scheduled a visit for 2nd August.

Tuesday, 7 June 2011

RCUK and HEFCE announcement to support Open Access


A majority of researchers would probably agree that Open Access is a positive way to disseminate research and reach a wider audience and the wider audience would probably appreciate their right to view research once barriers are lifted.   Removing the cost or avoiding the sale of a product is and has been shown to increase accessibility – take the aerial shots on Google Earth for example, how many home computer users have not snatched a look at their back yard or favourite holiday destination? The process is the same for documentation.
We sat down and discussed the number of ways that articles are being made and sourced as Open Access.  Our different backgrounds unearthed that we are often unaware of particular systems and procedures in disciplines outside our own fields and operations in other institutions. This serves to illustrate how complicated and mushrooming the idea of Open Access is, and how limited forms of Open Access have existed for a number of years.  Some kind of structuring and linkage between Open Access concepts appears a good way forward so the news of support from RCUK and HEFCE is welcome.
There are slight concerns about the clash of the peer review and so called pay to publish options (we realise these are the extremes and some hybrid versions are in place), as these can compromise and conflict with mounting pressures in the academic world.  Academics are aiming for journals with high impact (to meet REF needs), the ones outside the public domain that often come with the subscription.  The REF pressures far outweigh the need to dose every man on the street with detailed research findings.  On the other hand allowing Open Access to research and project documentation is an alternative opportunity to champion and publicise the achievements to similar and interested academic audiences whilst the general public can cast their eye over our achievements if they choose. Which system should we allow ourselves to gravitate to? 
How does this influence the HALOGEN group?  Much of our documentation is ‘white paper’ information on what and how we do things, information that we are willing (and proud!) to present to an open audience which also serves as a publicity agent and in effect enhances the purpose of our work.  In fact much of our documentation is (or will be) available, we are encouraged to contribute to our University repository and we have our HALOGEN project website.  The website is a source of information at different levels; short summaries for those with a passing interest and then the links and downloads provide detail to the audience who need to be more interactive or choose to know more depth in what we do.  There is an opportunity with Open Access to contribute to a bigger more widely accessible repository – but guidance is essential and we wait to hear the outcomes of this recent announcement from RCUK and HEFCE.
Andrew, Olly and Dave.

Wednesday, 11 May 2011

Progress Update for April

Good news - the resource issue that has held up progress in March and April has been resolved.
Olly Butters has joined the team and made an immediate difference.

Progress is summarised below:

Data Sources
 
The new Cappelli date source is now loaded into the database and being tested.

The 1881 Surname Census data has been reviewed and work to calculate parish centroids is well advanced. Once we have parish centoid data then the information can be added to the database and tested.

Tools for Researchers

The requirements for the 'web based query tool' have been reviewed and agreed with Jayne Carroll of Nottingham University's Institute of Name Studies (the owners of the Key to English Place Names data used by HALOGEN).

Work on developing prototypes has started and the target is to demonstrate the first cut products to the project team and key research users in late May/early June.

The first Project Board meeting is now scheduled for 27th May. Onwards !

Friday, 8 April 2011

Excitement at the GIS tutorial

Last week I gave a GIS tutorial using some prototype data we cleaned up from the original KEPN, PAS and GUL data sets to Turi, Jayne, Phillip, Dave and Mark W. The object was to help empower these blokes by showing them how to load up the data into a GIS environment and chop up the data with some simple querying methods thus stimulating the construction of new research questions.

Talk about the wow factor - they were really chuffed to see a spatial plot of what they used to know as rows and rows of tabulated data.  After filtering the data suddenly their hypotheses were mapped out in front of them, e.g. Place names with Cornish elements did gravitate to the county of Cornwall, place names with Norse language elements did gravitate to the North and East of England.  When ancillary data such as roads and rivers were plot as background layers I think the cogs and wheels started spinning and ways to answer research questions were suddenly looking so much easier for the researchers.

The data was questioned though, and quite rightly so, it should be a standard procedure for any researcher to be sure of the origins and quality of their data.

(i) Some of the grid references were slipping through our padding procedure and looking too accurate (by this I mean our rounding up of grid references to 0.5km). We did this to ensure privacy of data and maintain a consistent resolution between datasets. This is a small technical issue we need to address.
(ii) Cornish place name elements were detected in Herefordshire and way up in Lancashire.  In retrospect Dave and I examined the original data source a few days later and found that these results were true.  It was the original data that was throwing up the anomalies, technically the HALOGEN team appeared to get things right.

What do we learn from this? Firstly all the hard work is paying off and the researchers find this a really useful tool. Secondly we can only deal with the data we receive. We did our own quality check to be sure we had it right, if the source data is wrong HALOGEN cannot 'make up' data that fits, a strategy of quality control on the original data is required.

Andrew