Skip to content


Dear Edinburgh Fringe Website

Dear Fringe Website,

Your terms and conditions make me sad.

I work with the Web Science Trust and some of the big names in the Semantic Web and I was hoping I would be able to create “linked data” for the fringe festival. Linked data is the technique being used to publish government data on data.gov.uk and, according to Sir Tim Berners Lee, is the future of the web.

If I was able to do this (which I would happily do for free and with no bother to you), it would result in dozens of websites and phone apps remixing the fringe guide. While I’m sure your own iPhone app will be good (although I have a android phone, so no use to me), it would have been exciting to have 100’s of people providing alternate ways to work with the programme, and far more in the spirit of the fringe. Sadly it looks like the rules have been written from the perspective of advertising revenue and control, rather than fostering creativity and experimentation.

The Fringe will be awesome without linked data, but it could be and should be awesomer.

– Christopher Gutteridge.

ps. You really should rethink the policy “About linking by hypertext to our website” as it is unrealistic and draconian. I broke the terms and conditions by mentioning your URL in an unauthorised tweet.

Posted in Uncategorized.


Quick and Dirty RDF Reader

I know the world has quite a few RDF readers already, but I wanted something which was friendly and worked the way I do. I built the Q&D-RDF browser for my own benefit, but in the last few days it’s proved so useful I have made the effort to tidy it up a bit and announce it as a service.

It’s built on top of the excellent ARC2 library, and my graphite library. I’ve downloaded all the current prefixes from prefix.cc so that it will use short names for namespaces wherever possible. As it’s built on ARC2 it will handle lots of formats including RDF+XML, N3 and friends, and also RDFa and other abominations.

What I’ve found it most useful for is to point it at some RDF I’ve just completed and do a visual check that things look right. I’ve caught several dumb typos this way. The graphite “dump” format also seems to be much less intimidating for people that n3 or XML formats.

Dubious Data

“What RDF have you been writing?”, is what I’m sure you’re now wondering.

What I’ve been doing is brushing up my skills by writing some well engineered but deliberately pointless or inaccurate RDF datasets. However, by accident, one of them is quite useful. The playing cards dataset at http://data.totl.net/playingcards/  [Raw RDF, Browse] actually looks like it might be a good example to use in an introduction to RDF.

Posted in Graphite, RDF.


Using a triplestore instead of MySQL as a backend

I’m still looking at the barriers to using an RDF triple store as the back-end for a website. I’ve discussed some of this back in February already, but the problems remain unsolved.

Our usual pattern, when designing a website, is to identify the various types of entity that will be described by pages on the site. For an academic site we have some of people, groups, projects, publications, events, articles. We then create a database table or tables for each of these and php wrapper functions to get individual records, lists of records and methods to create & update records of each type. In PHP, we have an object representing the set of items (eg. Events) and an object representing each item. The SQL is kept abstracted away as much as possible.

The PHP classes which represent an item or a list of items, has methods for mapping the data into various formats; short HTML summary, an HTML page, RDF, XML, .ics, rss, atom etc. Occasionally some fields may be not shown to the public, for example if we use the same database for some internal administration.

On some sites, we have a table which stores all revisions of each item, and a table which maps each primary_item_id to its revision_id. Previous versions should never, ever be shown to the public as they may have contained errors or information we actively do not want to be public.

What I’m interested in is how normal web developers, rather than researchers, can achieve this.

I am still imagining a system with “classes” of things, like people and events, where the PHP is configured in such a way to be able to create/retrieve/update/delete individual “records”, that each triple will belong to only one record, and that we’ll have PHP functions which retrieve data from a set of records (by abstracting SPARQL instead of SQL)

Unanswered questions:

  • Internally, do we use our own namespace for the predicates or established namespaces (FOAF, SIOC etc) or a mixure?
  • If we use our own namespace, do we map into common schemas (FOAF,SIOC…) for the public view of .rdf data? Do we map it on demand, or when a record is updated? Do we expose our internal namespace predicates? I don’t believe just providing a mapping and let people map it themselves is a reasonable option.
  • Do we expose all of the triples? (what about ones used for administration? do we just make sure we have no secrets in the triplestore?) If so, how do we handle revisions? Have 2 triplestores — One for the public and one for admin? Or can triplestore SPARQL endpoints be configured in fancy ways?
  • How do we generate brief, unique URIs for items when they are created? In my experience URIs built from any of the meaningful data in an item are a mistake, eg. surnames etc. Using uuid’s are not an option — they are ugly. http://webscience.org/person/6 is better. My previous post suggested some solutions, and Talis have a weird solution using a pool of available IDs, but I don’t regard it as a solved problem. Then again there’s no standard solution in SQL databases.
  • If using tools to add value by importing/generating additional triples, how do we manage these? For example do we need to erase any of these if the records they refer to are removed or updated?

I think there are probably answers to all of these, but they need to be moved from ‘research’ to ‘development’. I’ll post updates if people solve any of these for me.

Posted in Best Practice, RDF.


MediaWiki Authentication Using Twitter and OAuth

The Dev8D wiki I set up for a recent JISC event uses OAuth to allow people to log in to the wiki using their twitter accounts (or users can register for wiki accounts in the usual way).

As promised in an earlier post, here’s a rough guide to how it was done.

1. Set up MediaWiki

I won’t go into details of how to do this here, but first step should be download and install a recent release of MediaWiki.

For a new wiki, I’d recommend installation of the reCAPTCHA plugin, to prevent automatic account registrations from spam bots.

I’d also prevent anonymous editing/creation of pages on the wiki, by adding the following lines to the bottom of your LocalSettings.php:

# Disable anonymous editing and page creation
$wgGroupPermissions['*']['edit'] = false;
$wgGroupPermissions['*']['create'] = false;

2. Create a new table in the MediaWiki database

Create a table named ‘twitter_users’ in your wiki database, with the following fields:

CREATE TABLE IF NOT EXISTS `twitter_users` (
    `user_id` int(10) unsigned NOT NULL,
    `twitter_id` varchar(255) NOT NULL,
    PRIMARY KEY  (`user_id`),
    UNIQUE KEY `twitter_id` (`twitter_id`)
) ENGINE=InnoDB DEFAULT CHARSET=latin1;

Note: If you’re using a prefix for your wiki database tables, this ‘twitter_users’ table will also need the prefix.

This table maps MediaWiki user accounts to twitter user accounts. It’s used to track whether an account on MediaWiki was created using twitter OAuth or not, and ensures only accounts created from twitter can be authenticated against twitter.

Without this, someone could create a twitter account with the same username as a non-twitter based wiki account (such as an admin account), and gain access.

3. Register a new twitter application

Go to http://twitter.com/oauth_clients, and follow the “Register a new application” link.

Fill in the fields as follows:

  • Application Icon: anything you like
  • Application Name: anything you like
  • Description: anything you like
  • Application Website: http://[your wiki base URL]/
  • Organization: anything you like
  • Website: anything you like
  • Application Type: Browser
  • Callback URL: http://[your wiki base URL]/oauth/callback.php
  • Default Access type: Read-only
  • Use Twitter for login: Yes

After submitting the form, you should get a Consumer key, Consumer secret, Request token URL, Access token URL, Authorize URL (make a note of these, or keep the window open somewhere for now).

4. Setup PHP OAuth library

I used the twitteroauth library for this (.tgz download).

This library requires PHP’s cURL library to be installed (package php5-curl on Ubuntu or other Debian-like systems).

Untar and unzip this into your MediaWiki extensions directory, and rename the directory to ‘oauth’:
cd /[wiki root directory]/extensions
wget http://github.com/abraham/twitteroauth/tarball/0.2.0-beta3
tar xzf abraham-twitteroauth-76446fa.tar.gz
mv abraham-twitteroauth-76446fa oauth

Recommended: Some of the code from this library needs to be accessible from a browser, so I’d recommend symlinking to this directory from the wiki root:
cd /[wiki root directory]/
ln -s extensions/oauth

You don’t have to do this, but it looks a bit neater than having URLs containing your wiki extensions directory.

Edit the config file in the oauth directory:
vi /[wiki root directory]/extensions/oauth/config.php
Set the ‘CONSUMER_KEY’ and ‘CONSUMER_SECRET’ to the values you got when you registered your OAuth application with twitter.
Set the ‘OAUTH_CALLBACK’ to ‘http://[your wiki base URL]/oauth/callback.php’.

To test that everything’s worked so far, visit:
http://[your wiki base URL]/oauth/
and click the button to sign in using twitter.

You should then be taken to a page on twitter.com which asks about allowing the application access to your twitter account. Clicking on the ‘Allow’ button should then redirect you back to:
http://[your wiki base URL]/oauth/index.php

Refresh the page, and you should see all the information twitter has passed back to the application.

5. Set up the wiki to use OAuth

Download TwitterAuth.php, and put it into the extensions directory:
cd /[wiki root directory]/extensions
wget http://github.com/davechallis/misc-scripts/raw/master/TwitterAuth.php

Modify your LocalSettings.php, and add the following lines:

require_once("$IP/extensions/TwitterAuth.php");

global $wgHooks;
$wgHooks['UserLoadFromSession'][] = 'twitter_auth';
$wgHooks['UserLogoutComplete'][] = 'twitter_logout';

Once you’ve added this, and signing in using OAuth worked as in the section above, try navigating to any wiki page. You should now be logged with your twitter username.

6. Additional Setup

Two last things need adding before we’re done:

6.1 Add a login button to the login page

Make a copy of the original, and then edit:
/[wiki root directory]/includes/templates/Userlogin.php

After the line which reads:

<p id="userloginlink"><?php $this->html('link') ?></p>

add the following lines:

<?php
$return = '';
if (isset($_GET['returnto'])) {
 $return = "?returnto={$_GET['returnto']}";
}
?>
<p>Or: <a href="http://[wiki base URL]/oauth/redirect.php<?php echo $return;?>">
<img src="/Sign-in-with-Twitter-lighter.png" alt="Sign in with Twitter" /></a></p>

Change the text/image above to anything suitable for your site (twitter has some preferred button images for this).

6.2 Redirect to the correct page after login

Some code needs adding/tweaking so that a user returns to the page they were on after logging in (the code added above for the login button helps with this).

Modify:
/[wiki root directory]/extensions/oauth/callback.php
and change the line near the bottom from:
header('Location: ./index.php');
to:
header('Location: http://[wiki base URL]/index.php/' . $_SESSION['returnto']);

And finally modify:
/[wiki root directory]/extensions/oauth/redirect.php
Underneath the line which reads:
case 200:
add the following:

    if (isset($_GET['returnto'])) {
        $_SESSION['returnto'] = $_GET['returnto'];
    }
    else {
        $_SESSION['returnto'] = '/';
    }

That’s mostly it!  I’ve probably forgotten a few things, and a lot of changes were made at the last minute/during Dev8D, so any fixes/suggestions are welcome.

Posted in Uncategorized.


Linked data vs Open Data

You can have data which is in a nice “open” format (eg. RDF/XML rather than HTML) but is not public.

Mild example; a FOAF profile for a member of staff should not be made public on the web without their permission, but could reasonably be made available to all members of the university.

Extreme example; exam transcripts should be made available in an electronic & machine readable form if a student (or tutor) wants them, but should be locked down.

You never want your students & staff giving their username/password to 3rd party apps willy nilly… although I bet they type it into SSH/Firefox/VPN on cybercafes so that ship pretty much sailed, the only difference is a smart phone app can be more targeted.

Here’s what I’m thinking a first draft of private linked data policy should be:

There is a web-page on the Intranet, requiring authentication which allows members to generate pass keys for their data, these keys will have a text description for what they are used for and a mandatory expiry date. They can then add “privs” to that key, like “profile data”, “timtable” or “assignments” (with carefully worded warnings about implications). They then cut and paste that key into the app they want to have access to the data.

The key consists of a username & password for the app to pass to the RDF server to get results. For example cjg-key2/8347e084309bc20 (where the username is a combo of my username and the key ID)

28, 14 and 0 days before the key expires the person gets an email telling them it’s expiring/has expired with a URL to click to add 12 months to the key. Maybe less time if it’s got very sensitive access.

RDF documents created with the username & password would contain a triple

<> generatedFor <...etc.../account/cjg#access-key2> .

And maybe and rdfs:comment with a deliberate long code in which the org can scan Google to catch if data is leaking.

Unexpected emergent behaviour

I can see one really interesting potential issue with this… If it became standard I can imagine high-pressure parents or even the government of the country an overseas student came from to keep a check on their grades! This is an interesting privacy issue, where a student can’t be *forced* to tell the truth to 3rd parties about their grades until their final graduation (or lack thereof).

Linked data makes this easier, but the problem exists already. You could just as easily insist the student gave over their normal username/password for parental or government monitoring.

Personal data vs Data I can view

Toby, one of our student coursework helpdesk guys, made a good point to me; there’s an important distinction between data I can view and data about me. A good policy may be to allow people to freely generate keys to any information about themselves (with some advice on the screen), but a key which can access raw data about other people, such as your tutees marks, or the internal university phonebook (with all those DPA-restricted names & numbers in) should require some more formal process, even if it’s data you can view via a portal in HTML.

That more formal process could be a signature, training or just a bigger EULA style page for you to fail to read.

Posted in Best Practice, Intranet, RDF.


Graphite Updates

I’ve done some more work on my Graphite PHP library. It now works a bit more like jQuery, which is nice. The general design philosophy is to try to make it as easy as possible to start doing something interesting.

I really like the following bit of code:

print $graph->allOfType( "foaf:Person" )->get( "foaf:name" )->join( ", " ).".\n";

Which prints out the list of all the peoples names. No loops! I even wrote documentation for it!

Graphite Browser

On Friday I needed a quick and easy linked-data browser and most of the ones out there I find bothersome and complicated, and I wanted something for a teaching aid, so I wrote this:

Which is very quick & dirty but I’m quite liking it to just look inside RDF & RDFa documents. It just does a regexp so all the hyperlinks go to http://graphite.ecs.soton.ac.uk/browser/?uri=XXXX instead of XXXX.

Posted in Graphite, PHP.


SIOC

This blog now supports SIOC RDF: http://blog.soton.ac.uk/webteam/?sioc_type=site

That is all.

Posted in RDF, Uncategorized.


Corkboards and Mind Maps

This is a cute tool to allow you to create and share virtual corkboards. I’m not sure how we could use it, but I’d like to find an excuse.

Also this week I’ve been using a mindmap to manage my scary todo list. Doing it in a webpage means it’s available via home, work laptop etc.

Posted in Uncategorized.


Turning alerts into stories

It is of interest when someone else writes something about the topic of one of our websites. Even more relevant when they write about the site itself. For example; http://webscience.org/ or http://users.ecs.soton.ac.uk/wh/

We use google alerts to keep our web/communications staff informed about anything being said about us, good or bad. Now and then one of the things is valuable to record and point people at. Currently this is done by hand, and is easy enough but any time someone is doing a repetitive job which takes more minutes every week, you gotta stop and see if a small script can help. There’s a tipping point where it’s worth doing a few hours work to save a few minutes of regular work.

Google alerts provide an RSS feed, so what we need is a tool which will allow our communications manager to see the latest few matches for a given search “Wendy Hall” “Web Science” “(ECS or (Elelectronics Computer Science)) and Southampton” etc.

Each feed item would have a “capture” icon which would save the title/link/summary into a database of “in the news/blogs/web” for that subject, but allow us to edit it afterwards for clarity.

Another interesting tool is http://purifyr.com/ which removes all but the “meat” of a webpage. A de-templater tool, if you will.

Posted in Templates, web management.


7 Degrees of Attention

In twitter, you don’t have friends. You have followers and people you follow. These are very much one-way arrows, compared with the facebook approach. Just because I see your tweets does not mean you choose to read mine (@cgutteridge should you wish to).

Today someone named @stuartbrown tweeted “anyone out there got info / point to info on RDF support in EPrints 3.2.1?”. I might have met him in the past, but I certainly don’t follow him, but my friend @psychemedia does, and so retweeted “RT @stuartbrown: anyone out there got info / point to info on RDF support in EPrints 3.2.1? #dev8d @cgutteridge” which drew it my attention and got him an answer in less than 40 minutes.

There’s often been talk of 6 degrees of separation and Bacon numbers and such, but these are very flimsy connections. foaf:knows style connections. It just indicates some basic connection between the two people. What today’s communication required was a chain of attention. @stuartbrown < @psychemedia < @cgutteridge. I’m wondering how hard it would be to work out the shortest chain of attention between two twitter users. What’s the shortest number of RT’s required to get my text to be read by decision maker X?

It’s not as painful as it sounds as “follows” lists are always much shorter. @nathanfillion may have 0.5 MegaFollowers, but only follows 91 people. However 2 hops is still around 10,000 users so would probably start to hit the API limits.

Of course, you could just mention them and they (may) see it anyway, but I was thinking more about how far your voice is from influencing decision makers via a channel they pay attention to.

Posted in twitter.