Skip to content


Studying the MPs

Last night Tony Hirst was trying to work out the birth-places and universities of the current UK MPs.

Here’s what I managed to produce while hacking with a glass of wine & TV on:

I did some seat of the pants data munging so I thought it’s worth explaining the process I used:

List of MPs

For the graph of MP’s birth-decade I did earlier in the year, I used a dbpedia relationship which gave me a nice little subject ‘member of 2005 UK Parliament’. This time around I can’t find anything so easy for the 2010 Parliament.

Solution, I used http://en.wikipedia.org/wiki/MPs_elected_in_the_UK_general_election,_2010 and a dirty little script…

my @tables = split(/<table/, join(”,<>));
foreach my $tr ( split( /<tr/, $tables[3] ) )
{
my @td =  split( /<td/, $tr );
if( $td[5] =~ m!”/wiki/([^”]+)”! ) {    print “$1\n”; }
}
(sorry for the godawful formatting, having hassles pasting code into wordpress since our upgrade.)

This then got munged in a text editor to create a .ttl file (much nicer way to express RDF than XML, esp. when doing hacky scripts)

This gives me this: data, [Browse].

Later I did something almost identical to produce a file adding affiliations as a party label and as an icon red/blue/yellow/other.

This gives me this: data, [Browse].

In retrospect I could have done this in one go, but it was late. Note the raw data of this file is just of the N-Triples format which is really easy to create and easy to import as RDF.

Making the Map Data

I then wrote a quick PHP script using my own Graphite Library to turn this data into geocoded RDF. eg. each resource as rdfs:label, geo:lat, geo:long and also an icon predicate I made up for the day.

View code: http://graphite.ecs.soton.ac.uk/experiments/parlibirth/mpmunge.php

View output: http://graphite.ecs.soton.ac.uk/experiments/parlibirth/mpmunge.php

As it’s a one-shot, I’ve just hard wired the relationship as “http://dbpedia.org/ontology/birthPlace” and then I just grab the results using curl. eg.

curl http://graphite.ecs.soton.ac.uk/experiments/parlibirth/mpmunge.php > born.ttl

…and then modify to “almaMater” and repeat.

Making Maps

This gives me files which can be loaded into my GeoRDF2KML tool. This forwards directly to Google Maps as that will accept the URL of a KML file as an input parameter.

Full disclosure; I added a dirty late night hack to geo2kml to accept my icon predicate to allow you to change what icon appears so I can get the by-party-colour-codes. If anyone has a ‘proper’ predicate to relate a geolocation to a map icon, let me know and I’ll support it.

To make things simple, I used ‘curl’ again to save the KML files to the same website.

Final Maps:

Note that the data is patchy. It only shows MPs with a geocoded birthplace/university listed on dbpedia.

Posted in Geo, Graphite, Perl, PHP, RDF.

Tagged with .


New Tools

As well as the more experimental stuff, I’ve also produce several more useful tools:

sparql2kml

http://graphite.ecs.soton.ac.uk/sparql2kml/ – This takes a SPARQL query which returns ?lat,?long (or ?georss) and ?title and maybe ?desc and ?placename and produces a KML file so you can see it on Google Maps or Earth!

As an experiment I used the following to find the birth place of Southampton football players.

PREFIX dbo: <http://dbpedia.org/ontology/>
PREFIX dbpedia: <http://dbpedia.org/resource/>
SELECT DISTINCT ?georss ?title ?placename WHERE {
?person dbo:team <http://dbpedia.org/resource/Southampton_F.C.> .
?person dbo:birthPlace ?place .
?place <http://www.georss.org/georss/point> ?georss .
?person rdfs:label ?title . FILTER langMatches( lang(?title), “EN” ) .
OPTIONAL { ?place rdfs:label ?placename . FILTER langMatches( lang(?placename), “EN” ) }
OPTIONAL { ?x <http://dbpedia.org/property/county> ?place }
FILTER (!bound(?x) && ?place != <http://dbpedia.org/resource/England> && ?place != &l
t;http://dbpedia.org/resource/Wales> )
}

 

View it: Google Maps or KML for Google Earth.

excel2csv

This one is dead simple. It converts an excel file into comma separated values.

http://graphite.ecs.soton.ac.uk/excel2csv?src=http://opendata.s3.amazonaws.com/bridge-weight-limits-2010.xls

sparqllib.php

http://graphite.ecs.soton.ac.uk/sparqllib/

Nice and simple library to let you use SPARQL from PHP. The function names are deliberately copied from the mysql ones so you have sparql_connect, sparql_fetchrow etc.

 

Posted in Uncategorized.


Linked Data Experiments

So I’ve been doing lots of little experiements with consuming open data…

 

 

Posted in Uncategorized.


Blue Plaque

So I’ve just been to the Open Data Hack Day in Oxford which was good fun. Met some cool people, wrote a lot of code and drank some brandy.

My team was playing around with using dbpedia‘s data mixed with geo-location to find you an interesting fact about where you currently are. We had a lot of fun with it — the final results are here:

It does some neat things. It uses javascript to ask your browser where you are, or failing that to use the wikipedia name of a city, the lat/long or use the postcode. http://data.ordnancesurvey.co.uk/id/postcodeunit/SO171BJ will give you the lat & long thanks to @gothwin.

It then attmpets to find nearby places on wikipedia which are the hometown of something. It does this by searching for things within + or – 0.2 of a latitude and longitude (I know that’s not going to be a perfect square, but meh). If it finds nothing it doubles the search range and tries again until it does.

It then gets all the things that have the city as a hometown, picks one and renders a blue plaque.

For added sillyness, if there’s a image available, it has a little proxy which downloads the image, shrinks it to no more than 300×300 to be phone-friendly, and makes it white-on-blue to match the plaque.

I stole the style of  buttons at the bottom of the page from m.ox.ac.uk which is an excellent example of how to make a website to work on a phone, rather than bothing making a specific phone app.

We won ‘most creative use of data’. Some of the other groups did more worthy things like visualise arts-funding data and make useful bus timetables so forth. One group had a great idea but didn’t get very far which was linked-data top-trumps. Each site in the linked data cloud has quite a few stats so you could probably do something cool with that. Most triples, most links to other datasets, most open license… Actually I wonder if there’s a tool out there which you can feed a csv and it’ll produce you nice pdfs of top-trump cards to print-out.

Posted in Uncategorized.


UNIT4 and Linked Data

I was recently in London for a meeting with some folks from UNIT4 about their recent forays into linked data with their Agresso Business World software.

They have the advantage of having a large installed base (>90 local councils, and >250 Futher and Higher Education institutions in the UK), so can hopefully provide a mass of data without customers having to set up or install additional systems/infrastructure.

They’ve initially been looking at local council data, with a view to widening this later (Universities are an obvious choice, especially with the growing interest and deployment of institutional open data).

Local Councils

Local councils will have to comply with the Prime Minister’s call to publish financial transactions over £500 from January 2011. Being able to do this simply, with an existing system (which already holds their financial information), makes a lot of sense to the council, while providing the data to the community in an open way.

The Guardian‘s Data Blog has a great summary of how this has been done to date: Local council spending over £500: full list of who has published what so far.

The whole list ranges from one to three stars of the Linked Open Data star scheme. Being one of the first councils to move up to four or five stars certainly couldn’t hurt…

UNIT4 are currently running a pilot with the borough of Windsor and Maidenhead, who already make a lot of their data open (1-3* Excel/CSV/PDF mostly). UNIT4’s plan is to take them up to 5* data, with a view to using the same techniques, software and lessons learned to do the same for other council.

From 3* to 5* Data

They’ve been looking at workflows for converting from existing financial data to RDF using the Payments Ontology, aiming to generalise to the process so that the same software and techniques can be applied to non-financial data an organisation might have.

Other ontologies used include VoiD and RDF Data Cube.

Redaction is obviously an important feature here, which it seems Agresso supports natively. The Payments Ontology also supports redaction (and I think there’s also an extension to it which supports redaction in a more fully featured way). This is something which can’t easily be automated though, and will still require human effort to clean up data before it gets opened.

This a great way to get a foot in the door – having one successful workflow from CSV/XLS to RDF means that an organisation can easily apply it to others, with the same software and input formats. Though this is an area that I’m guessing a lot of software providers will want be the center of…

Work done so far for Windsor and Maidenhead can be seen here: Local Government Spend Explorer

The hard parts…

The meeting also raised some familiar concerns/questions about the publishing and maintenance of open linked data:

How do organisations agree on identifiers to use for suppliers?

This is pretty hard without a central registry or lookup service. Companies House data would be a great starting point, but is not open or free.

UNIT4 are going down the route suggested by Tim BL – minting their own URIs for things, then using the owl:sameAs predicate to link them to definitive versions later.

How should an individual entering data find out an supplier’s URI given its name?

Auto-completion? Drop down lists? Even though this is more of a user interface issue, it raises the important point of getting people who don’t cared about linked data to be accurate about data they’re entering.

Which URIs should we use to describe currency?

While ISO maintains a list of currency codes, they’re not available in an open form, and the data set isn’t available without paying.

How should data from separate councils be aggregated?

There are hundreds of local councils in the UK, and collecting data from all of them, or querying 100+ triplestores to get at data for comparison just isn’t feasible.

I’m guessing this is something we’ll eventually have to face in the Higher Education linked data world (e.g. someone wanting to query Universities for course data won’t want to have to connect to download data from dozens of institutions).

Should there be a central registry? Should data.gov.uk pull local council data into a central triplestore? Should UNIT4 be pulling in the data as a service to their customers? Are any/all of these methods sustainable?

What next?

I’m sure we’ll be hearing more about UNIT4 and linked data in the near future (assuming the Windsor and Maidenhead pilot goes well!). If the strategy and data produced is successful, we may well see a number of councils adopt it.

If this happens, this would be a great starting point for producing institutional open financial data – choices of identifiers and ontologies to be use will be much clearer if there’s a large body of homogeneous data out there.

Posted in Uncategorized.


Barcamp Southampton

Last weekend I helped run the first ever Barcamp Southampton.

As this isn’t actually part of my day job,  the write-up is on the SoTech blog, instead.

Posted in Events.


Notes on SITS – the Scholarly Infrastructure Technical Summit

I was recently sent to attend the Scholarly Infrastructure Technical Summit (SITS) by JISC, along with Ian Mulvany from Mendeley.

The goal of each SITS meeting (as I see it) is to get a technical experts (a mix of developers, and project managers who understand tech) together to talk about their experiences/problems/successes with various scholarly infrastructure tools or components.

Something which worked extremely well (which was new to me) was having the meeting run according to an Open Agenda. Topics of interest were brought up by participants, and then voted upon to ensure that there was sufficient interest in a specific subject area.

The meeting started with a run through of topics that were raised at the previous SITS. These included:
SWORD, reverse SWORD, common tooling for workflows, storage abstraction layers for repositories and author identification, as well as discussions around appropriate citations for digital objects.

The topics which were raised and eventually discussed this time around were:

  1. Authentication/Authorisation
  2. RDF/Linked Data
  3. Web Archiving
  4. People/Author Identifiers
  5. Microservices
  6. Curriculum/Training Development
  7. Search Engine Optimisation (SEO)
  8. Lightweight Languages

Authentication/Authorisation

Discussions here were focused around authorisation and authentication, both across services within an institution, and between institutions.

Shibboleth was (as expected) the primary technology talked about, followed by OAuth, and then OpenID.

I was hoping to hear of some success stories here, but mostly it was the problems and questions about these systems which came through:

  • How can we combine Shib with IP or key based auth?
  • How can we provision temporary/guest accounts in such a system?
  • How can we trust remote credentials?
  • How can services authenticate between each other?

The issues with service to service auth were mostly based around the fact they’d require extra client libraries to be used (especially for web services, where basic access authentication remains the easiest method to use).

Another major issue was that of management of access to resources. If centrally managed, how can a data/service provider be confident that their institution is keeping their groups/users/access levels up to date and correct? And more importantly, how can we be confident that an external organisation who has access to our systems is doing the same?

Related to this issue is that of spreading access control of data around (Chris Gutteridge also has a blog post related to this in linked data). Most major auth systems seem to centralise control – but what happens when an individual wishes to share some of their personal data with another person/group/service? What happens if a department wishes to have their own policy controls? What models are there for delegating control, whilst ensuring overall stability and security of a system as a whole?

I think the most interesting outcome from this topic was discussion about a shim or meta-auth layer that could sit behind several auth systems. There certainly seemed to a be a lot of interest in something which could authenticate against Shib/OAuth/OpenID/anything else, and then provide a single set of auth details to an institutional system.

It would mean that institutional software could interface with this one layer, and have additional auth mechanisms added to it through extensions/plugins, rather than having to plug multiple auth systems directly in, and have to update code each time a new auth system comes along.

Mendeley do this internally, but if there’s an open source solution for this out there we don’t know about, I’m sure it would be very popular…

RDF/Linked Data

The discussions around linked data were very similar to several I’d had with people over here in the UK – the barriers to adoption were the same at least:

  • Which vocabularies should we be using?
  • Should we be creating our own?
  • Where can we see examples of best practice?
  • Whose identifiers should we use?
  • What content should we be making available first?

Seems it would be good to include some US institutions in the discussions that are happening around the UK academic community at the moment, at the very least to prevent us from going off in entirely different directions!

Lack of tooling was also perceived as a significant barrier to entry. There was a lot of interest in access to linked data using RESTful APIs (e.g. the linked-data-api), and using javascript and JSON to consume the data. These allow experienced developers to consume RDF using methods/technologies they’re already familiar with.

During this discussion (and during a couple of others), quite a few people expressed dissatisfaction with Dublin Core as a means of describing repository data. It seems that some (including people at Google) were interested in looking at HighWire as an alternative. I know next to nothing about this though (a search doesn’t reveal much either…), but will update should I find out more.

Web Archiving

Next up was the topic of archiving web resources, which followed on nicely from a presentation on recent developments in Memento at the DLF Fall Forum.

A key point here was the issue of when/what to archive. Some web resources (a paper in a repository for example) have fairly well defined versions – but do we want to archive every single one?

Other resources (a page aggregating 3rd party content for example) won’t have such well defined versions, and will need to either be archived periodically, or by constantly watching them for changes.

Assuming we’ve taken care of the above, the next point of discussion was about searching. What sort of interface will be needed to search historical resources? Obviously a user won’t want to be presented with a dozen search results containing almost identical content. They’re also likely to want to browse back and forth through important versions of an item onc they’re found a document they’re after.

Someone also raised the point of the difference between browsing/searching historic documents vs. browsing/searching as if you were on a historical system. Some users will want to search using historical indexes, others will just want historical results.

There was certainly interest in Memento as an easy to implement strategy for archiving some web content now (adding it to repositories would be an easy win in this regard), and then worrying about what to do with the data later. It also seems to have the advantage of working below the normal web application level, meaning that the same technology can be used for archiving video/images/RDF/html, without requiring an application specific setup each time.

It was also mentioned that JISC would be commissioning some large scale work on the preservation of fast moving resources.

People/Author Identifiers

The focus of this topic touched on a lot of things, but mostly revolved around ORCID (Open Researcher & Contributor ID).

The ORCID initiative aims to provide a registry of authors/contributors (to aid in communication, author disambiguation), which can then be linked to other ID schemes, to publications, or to each other.

ORCID is (I believe) a follow-on from Thompson Reuters’ ResearcherID. ResearcherID required self registration though, which is where it’s believed to have failed (they had <20000 individuals register). ORCID’s aim is to get author information form institutions, rather than individuals.

I was surprised to see there had been so much interest in this already, 300-400 institutions have already registered their interest in it. It seems that some journals may start requiring ORCID IDs before publishing, which could well be a driver in this.

This, along with the fact that it should work nicely with other ID systems makes it look like something worth keeping an eye on.

It’s not without its issues and potential problems though.

The first of these is that the information kept by ORCID hasn’t been finalised yet. What should they store along with an author’s ID and name(s)? Publication list? Grants? If so, who’s going be responsible for maintaining it?

This also raises the issue of control of personal data. If an institution makes a statement about you in ORCID, do you have the right to retract it? What about if an individual starts making statements an institutions knows are untrue?

Storing the provenance of each fact about an individual in ORCHID seemed to be the accepted solution for this – it would leave it up the data consumer to trust individually/institutionally submitted facts about a person.

By far the biggest obstacle seems to be the lack of an ongoing business model for ORCID though. Once it’s up and running, and as it’s a centralised service, how should it be funded? The identifiers it mints will need to be permanently resolvable in order for anyone to trust it and use it as a service, so how can the community guarantee this?

The project is still very much looking for contributors to develop things from an institutional side though, looks like there may be several potential projects here.

Microservices

This is a term I hadn’t heard before this month, but came up a lot at the DLF Forum, as well as being a topic with significant interest at the SITS meeting.

Luckily it wasn’t just me who was unsure of exactly what it meant – it seems to be a bit of a buzzword, taken to mean different things by different people. A rough summary can be found on the iRODS micro-services page. The CDL (California Digital Library) also have their own take on microservices on their Curation Micro-Services pages.

The basic gist of microservices seems to be this: Repository software is made up of collections of services. So why not separate them out, and make each one available for reuse, either as a web service (SOAP/REST), or on the command line?

This mirrors the Unix philosophy of making programs do one thing well, and making others by combining several of these.

They allow for easily interchangeable components with a narrow focus, allowing for complex services to be built up without reinventing the wheel each time (especially when moving between languages or platforms).

Some examples of microservices might include:

  • image resizing
  • file checksum calculation
  • file hashing service
  • send email
  • object storage

I’m not convinced the specs for these are well defined enough for general purpose use yet, but I can see the technique in general being very useful. If the same microservices can be called through multiple interfaces (command line, REST, etc.), then it should in theory make them language agnostic.

Curriculum/Training Development

Things moved to a slightly less technical theme here, the focus being on how to get new staff/project members quickly up to speed in a development or project environment.

The key point was to work out how best to ensure that new team/project members gain the skills they need to get their work done.

This includes a mix of technical and non-technical skills, the exact nature of which will vary depending on the project:

  • Source code management (git/svn) and committing guidelines
  • Documentation guides (how to document code/software)
  • Code style guides (what should my functions be named? should I indent with spaces or tabs?)
  • Unit testing (which framework should I use? when should I write them?)

There are also of project/platform specific things an individual will need to know:

  • How do I add plugins to this repository?
  • How do I code X in language Y?

A key point raised here was in the difference between people who were primarily librarians and those who were primarily developers. How do we get each up to speed in the areas they lack? A single curriculum for everyone probably wouldn’t suffice.

Additionally, how should this be taught, and who by? Online notes? Or taught as part of a library/computer science course?

Making sure that the right people attend the right training courses/workshops was also a key point mentioned. Ensuring that an individual has the prerequisites necessary to participate is essential in order not to waste time and money.

Some suggestions about starting points for developers/managers looking into this included the following:

Search Engine Optimisation

This topic should really have been called “Google Scholar Optimisation”, such is the perceived importance of Scholar in the library/repository world (Scholar is apparently way ahead of Web of Science as a student portal to research for example).

There was a great deal of dissatisfaction expressed with way Scholar works, primarily concerning the fact that it’s the Scholar team who are dictating the metadata that institutions produce in order to be included in their index (more info in Google Scholar’s Inclusion Guidelines).

It was largely felt throughout the room that it should be the academic community who is responsible for agreeing upon a standard for exposing metadata (RDFa was mentioned here), rather than being forced to adopt 3rd party’s schema which doesn’t fit their data very well.

One thing I learnt from this was that the Scholar team is separate from the regular Google search team, and the technologies they use to index/harvest differ. This means that institutions have to produce metadata multiple times: for Scholar, for regular google search (and probably more for additional harvesters).

So what are the solutions/workarounds? Some ideas raised were to:

  • Agree on RDF(a) standards for presenting metadata
  • Contact NISO about developing a standard
  • Approach a rival to Scholar (e.g. Microsoft Academic Search)
  • Use the collective bargaining power of institutions to effect a change at Scholar

I’ll finish by saying that quite a lot of people in the room felt very strongly about this, it didn’t sound like Scholar was making anyone very happy!

Lightweight Languages

The main gist of this topic was discussion about barriers to introducing a new language into a project or environment.

Ruby was the main language being discussed, but the discussion could just as easily apply to any language/technology being considered for use.

The key word here was “misconceptions”, seemingly from all sides where introduction of a new technology is concerned.

One barrier to adoption was seen to come from developers themselves. Many are reluctant to learn a new language (especially one perceived as too hard/different to the ones they know). This could be especially relevant to non-CS developers – they’re more likely to have language specific experience, rather than abstract programming knowledge, making it harder to switch.

The next was from a sysadmin point of view. The introduction of new languages can be seen as a security risk, and as yet another set of software that needs patching/updating/configuring. Different languages also have very different security models that need looking at, PHP’s now deprecated safe mode and Java’s Security Technology are a couple of examples which are very different indeed.

There’s also a certain level of suspicion (speaking from my own experiences here too…) about whether new languages/technology are really required for a project. If a project manager has a good reason for picking a certain language (e.g. they have to use a specific software package, libraries/plugins in a certain language are good, a language is needed to easily interface with other tools, etc.), then that’s one thing. There are just as many less valid requests to use a language though (e.g. it’s all I know, it’s a current buzzword, etc.).

So how do we get around these issues?

Virtualisation or bundling of the language was one solution mooted. Using virtual machines is one example of how this could be done (though it still raises many questions about security and trust). The other example given was in bundling up the language and libraries in a single package that could run in a more sandboxed environment (the Ruby language bundled in a WAR file running on a JVM was given as a successful application of this technique).

More important than this though, was the idea of getting sysadmins involved early. Rather than going to them with your requirements, it seems that teams had much more success by involving them with developers from the start, getting them on board to discuss issues and solutions, rather than dictating them at a later date.

Closing Thoughts

Overall, I thought the meeting was a big success, and opened my eyes to quite a few big things that I wasn’t aware of before (ORCID, microservices and Google Scholar issues being foremost among them).

More importantly, each topic mentioned above was finished with some action items, so I’m hoping that we hear some progress on these from various SITS attendees in the near future (I’ll add links to this post as I hear about them).

Future SITS meetings will take place in different locations, and with different attendees, so I’m hoping we’ll see a good cross-section of issues and experiences coming from the meetings. I’m sure we’ll see a lot of common threads come up that lead to more less repetition of software development, and more importantly, less repetition of mistakes!

Dave Challis
dsc@ecs.soton.ac.uk

Posted in Uncategorized.

Tagged with , , , , .


Auto Discovery of voID via SPARQL

I wanted to know more about the data in a SPARQL endpoint. I had a good idea, search the endpoint for triples that mention the URL of the endpoint… no results.

I tried a few endpoints and most returned nothing, but trying one of the RKB Explorer endpoints I got a single triple. But it was the right triple!

From this I can disover everything I need. I suggest that this should be best practice; SPARQL endpoints should contain voID for the datasets they contain, and relate themselves to the datasets using void:sparqlEndpoint.
Even if the voID just has a human readable title and description that’s infinititely better than nothing.

Posted in Uncategorized.


Everybody needs a 303

So there’s a lot of debate the past few days about the issue of 303’s being one of the two accepted ways to get from a URI for a non-document (eg. the City of London) and for a document about that thing (london.rdf or london.html etc.)

One of the key points Ian Davis made is that it must be practical for people on crappy ISP setups (or who’s computing services don’t let them touch their server configuration etc. Ian Davis and Leigh Dodds have done some useful work suggesting workable patterns for people to follow, and I think what’s needed next is a ‘cookbook’ of how to achieve these patterns using available tech.

With this in mind I’ve been trying out solutions and would like to present a simple solution for how (nearly) everybody can have a 303. I’ve tried this on apache on local Redhat Enterprise and Ubuntu servers, and also on the data.totl.net server on Dreamhost — it works on those. It nearly worked on the apache setup which comes with OSX, except for the fact it redirected with a 200 rather than a 303.

What you do is put your RDF files in a directory; eg. /project/people/marvin.rdf and define the URI for marvin as /project/people/marvin

Put as many in as you like then add this .htaccess file:

RedirectMatch 303 (.*)/([^\./]+)$ $1/$2.rdf
<Files ~ ".rdf">
 ForceType application/rdf+xml
</Files>

The first line just redirects *anything* without a dot “.” in to the same address with .rdf on the end. This means that /foo will direct to /foo.rdf and then 404, but who cares? The second bit just makes sure that the mimetype is correct if the server doesn’t do the right thing with .rdf

There are better ways to do this, but this works on a pretty vanilla, and locked down, apache set up.

So there you go, one less excuse. Link that data!

Posted in Uncategorized.


Searching a SPARQL endpoint

Recently, OUseful Blog has been talking about how to get started hacking SPARQL queries. So here’s a simple one. It looks for things with a search string in their name, title or label:

PREFIX foaf: <http://xmlns.com/foaf/0.1/>
PREFIX dct: <http://purl.org/dc/terms/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT DISTINCT ?thing ?name ?type {
 { ?thing foaf:name ?name }
 UNION { ?thing rdfs:label ?name }
 UNION { ?thing dct:title ?name  }
 OPTIONAL { ?thing a ?type }
 FILTER (REGEX(?name, "YOUR-QUERY-STRING-HERE","i"))
}

For example, searching for Ventnor in the Ordnance Survey. I suspect it’s not that fast because it’s actually having to work through filtering a huge pile of data. I thought searching for “^Ventnor” might be faster (It would in SQL as the indexes can do string-starts-with quickly), but it doesn’t seem to be. Advice on optimising?

If people are interested, I could add this as an option to the Graphite SPARQL Browser.

SPARQL/SQL Translation

For SQL users, UNION is in effect an “OR”, OPTIONAL can be thought of as “LEFT JOIN” and FILTER as a WHERE.

If the SPARQL endpoint were an SQL database, it would be a single table containing three columns, subject, preficate and object. (Yes I’m skipping some stuff here to keep it simple). I’m going to remove the UNION for now as that’s basically like running several SELECTs and merging the results. Note that “a” is an alias for “rdf:type”.

SELECT DISTINCT ?thing ?name ?type {
 { ?thing foaf:name ?name }
 OPTIONAL { ?thing rdf:type ?type }
 FILTER (REGEX(?name, "YOUR-QUERY-STRING-HERE","i"))
}
SELECT DISTINCT t1.subject, t1.object, t2.object FROM
triples AS t1,
LEFT JOIN triples AS t2
ON t1.subject = t2.subject
WHERE t1.predicate = 'http://xmlns.com/foaf/0.1/name'
AND ( t2.predicate = 'http://www.w3.org/1999/02/22-rdf-syntax-ns#type' OR t2.predicate = NULL )
AND ( t1.object LIKE '%your string here%' )

Although I’ve changed the regexp to a LIKE. I’m not 100% sure I’ve got this entirely correct, but it should give an SQL hacker a feel for what’s going on. Every triple in the SPARQL select is effectively an inner join where the named parts ?foo are joined to the columns they were associated with in the previous triples. You can do some very funky things in SPARQL, but you need to get joins from lesson one. Even a trivial query  on a property of a field will probably require you to add a { ?item a “foaf:Person” } or you’ll get all things of all types which isn’t what you’re going to want.

I think that as RDF and the semantic web achieves escape velocity [PDF], we’ll need to make some tutorials for people who just want to get the job done. Right now we’re still working with almost entirely early adoptors. We need to make getting data out of SPARQL achievable for people who don’t really care. I found a PHP library for working with SPARQL, but it seems to be from more than 5 years ago. Perhaps I should write a SPARQL library which looks like an SQL library? sparql_query() sparql_connect() etc? (comment if it’s worth my time…)

Redirect to SPARQL

Dave Challis had an interesting suggestion yesterday… Making a URL which accepts a ?q=XXX query and redirects to the SPARQL query that searches relevant labels in our endpoint. That way we select which predicates we consider labels, and it gently cues people into SPARQL without forcing the initial learning curve.

Update:

@ldodds points out that the O.S. endpoint has some funky Talis features, so that there is a simple search API. Which gives pretty useful results. I’ve seen, passing, searches which return RSS, but what I’d not realised until today was that the RSS contains lots of useful triples, so in effect it’s just a structured list of RDF descriptions. This approach looks very useful for some usecases I’ve been thinking of. Specifically how to make it easy to search an organisation’s datasets. For example, how to find a building at southamton university when all you know is “Zepler”.

Posted in Uncategorized.