Monday, 22 March 2010

The Sprint Cycle


So that I don't "bite off more than I can chew" -- my aims are to make many iterations of the Sprint Cycle (at a guess, four-ish). The dotted areas on the diagram above show the gaping holes in my knowledge. The first iteration will be about getting a better understanding of each of the areas and finding the simplest tools for working with them.

Sprint 1: The aim to to produce the simplest thing possible that is remotely useful. I will attempt (this week) to make a full traversal of the cycle. At this point I don't have the details of people's profile on social media sites (the survey hasn't been published yet), I don't have any real knowledge about worth with repositories, nor of working with LinkedData. I will need to work with what I have, namely the data that is already "out there"... web pages and links.

Whilst the aim may be to make something more like the Research Portal (shown below), this week is going to be all about doing something MUCH SIMPLER.




Aims

  • Review crawling components (and mine York's existing data - wherever it is)
  • Display content in a TagCloud
  • Attempt to integrate with some LinkedData
  • Attempt to present and data or relationships in a novel way
  • Publish the Social Media Survey to hopefully find out more about what is being used at York.
The first iteration of the cycle looks like this. And nothing like thrashing around in frogspawn.






Milestone 1 Report

I have trialled a number of tools (SocialText, Jive, Elgg, Confluence, LifeRay Social Office and Cyn.in) with a number of teams which include...


Thursday, 18 March 2010

Milestones

Milestone 1: People and Site

a. People. The first milestone is about assembling a team of people engaged with the as yet, non-existent PPPeople PPPowered technology. This team will provide the data, particularly their social media usage and profiles etc with which the initial data-mining can be built upon.

Goals: To have at least 20 people in a number of teams of people willing to work with me on the JISC project, providing hard data, usage and feedback.

b. Presentation. The data that is ultimately mined needs to be "hosted" in a software tool that provides a level of service that is valuable enough to be used frequently. This will ensure that meaningful relationships between people are presented and then subsequently pruned by participant.  The choice for which tool we use at this point was initially between Wordpress, Drupal, Elgg and LifeRay.

The conceptual framing of what this tool should be, a people directory, a people discovery tool or a profile repository or a personal new aggregator is quite important with regards to setting expectations (for re-visiting the site)

Outcomes: 

  1. To have a site up and running, available to staff at University of York that at least displays their profile in some way.
  2. A blog post reporting on the project so far, the plan, the tool chosen (with reasons) and invitations to help with the plan (which semantic repositorities to us (and how), which crawling tools to use
  3. Review the approaches take with other social networking sites to gather any useful approaches. For example, are their benefits to be gained by leveraging an existing social network and evangelising the use of a certain social network in order to ease data gathering.

This phase is very dependent on peoples' time availability and willingness to engage and also on the richness of the data returned from the survey. If nobody at all uses social media or completes the survey we won't have a lot to begin with.


Milestone 2: Just Data and Display




The first stages of displaying mined data with avoid the complexities of semantic reasoning about the data, about using unusual sources of data and begin by simply asking people to complete a survey. This will hopefully result in a spreadsheet of sites, blogs, social media memberships (twitter, linkedin, CiteULike etc) and will form the basis for exploring slightly less explicit data (connections in Linked in, followers on Twitter, mentions in Twitter or on other people's blogs.

To display this data, initially we will need to adapt the profile in the presentation tool, perhaps including an RSS aggregation.

Outcomes: 

  1. A survey asking people to identify their social media accounts, interests in the form of keyword, URLs etc 
  2. A site with at least 20 members showing their "simple" social media membership and some of the data (recent tweets, friends etc). 
  3. The initial adaptation code (module/plugin) uploaded to the SVN site.
  4. A blog post showing developments and with comments from members.

This will require a good understanding of the plug-in architecture of the presentation tool. The whole point of this project is to be integrated into a tool that is usable and used.


Milestone 3: More Deeply Mined Data

This stage looks to find information beyond that that is given, perhaps crawling Google searches, finding more distant links between resources. Ideally any crawling or querying tools should be well integrated into the presentation tool OR componentized and talk to the presentation tool via XMLRPC or similar.

We will experiment with (probably python-based) web-crawlers such as Domo, Harvestman, Mechanize and Scrapy. In addition we will look at free or low lost tools and services available. These may include...


  1. Yahoo Pipes - online data manipulation and routing
  2. 80 legs - a crawling service
  3. Maltego - a desktop open source forensics application
  4. Picalo - desktop forensics application (or any other tool)


Outcomes:

  1. More interesting data, richer connections
  2. A generic crawling methodology 

 Whilst I am more than familiar with creating simple crawlers this will need a more standalone, robust, better architected approach.


Milestone 4: Semantified Data

This stage will look at how to apply semantic tools, or a more reasoned understanding of the data gathered. We will be looking at ...

  1. Yahoo Boss for general searching.
  2. OpenCalais, to attempt to discover entities in unstructured data
  3. DBPedia, Edina, Freebase for understanding entities like Towns, Universities, Concepts etc
  4. ePrints research repository to connect people via research outputs
 SPARQL is very new to me. I need more understanding of RDF, LinkedData etc. Although the JISC Dev8D conference gave me new insights into the possibilities presented by LinkedData and open data I still feel I have a way to go to fully understand this area.

Milestone 5: Slightly More Sophisticated Presentation

Depending on the data gathered, there will be opportunities to present the data in more interesting and usable ways. Initially we will attempt to use simple visualisations such as timelines, tree maps, word clouds etc. which can be easily integrated in the presentation tool.



More ambitious visualisations such as Mention Map example will be explored if appropriate.

 Need to understand more about the maths behind networks and visualisation. Luckily Gustav Delius is around to advise.








Tuesday, 16 March 2010

Basic Plan


This shows the basic plan for the project. The hope is that the first run of this process should be complete as soon as is possible. This means that the REASONING and VISUALISATION stage may be omitted, where the the data can be presented as a simple list.

This makes the choice for the DATA-GATHERING and PRESENTATION tools very important. We want them to be very well integrated. Ironically, given the importance of the PRESENTATION stage, namely what user profile data it stores (social media usernames etc), what features it supports (RSS aggregation, Twitter presentation, Plugin development opportunities etc).

Saturday, 6 March 2010

Tuesday, 16 February 2010

Creating Teams of "Tyre-kickers"

Collaboration is impossible to fake


One of the problems with evaluating collaboration tools is that they can't be compared on features. One tool's blog can be very different from another tools blog. A small missing feature from a wiki can make it less easy to use.

The only real way to test which collaboration tool will best suit the PRESENTATION layer, the part where any mined data is displayed and used is to use the tool with a number of groups of real, live, collaborating people. With this in mind I have tried to engage a number of teams willing to "throw their all into testing an environment"... their nickname is "Tyre Kickers".

A large component of this projects success will not depend on whether or not the technology works but whether or not people use it.

One of the problems with this Tyre-Kicking approach is that once people have invested time and effort into working with, and more importantly around a tool, they may have grown to like the tool despite its warts and be loathe to move on to another tool because all their content exists somewhere else.

Aware that the different teams will have differing uptakes in terms of engagement, some will want and need to use the tools whereas others will dabble. Some groups will work in the same office whilst others will be dispersed, across and beyond the university.

The Tyre-Kickers I have held meetings with and introduced to up to three tools currently are....



  • Applications Deployment - A Computing Service team
  • Biology - Development Team
  • Collaborative Tools Project
  • Communications - Small departmental team
  • Computing Service - Large departmental team
  • Court, Country, City. British Art 1660 - 1730 - Humanities project team
  • Digital Library - Development team
  • Finance at University of York - Small department
  • Humanities Research Centre - Large cross disciplinary departments
  • IT Support Office - Small team
  • Liaison Librarians - Large dispersed team
  • Marketing Interest Group - Dispersed team
  • Mathematics - Small team
  • Planning at University of York - Small team
  • Research Themes - Small project team looking at strategy
  • Science and Technology Studies Unit
  • Sociology - Small team
  • Stockholm Environment Institute - Large collaboration team
  • Sustainability at University of York - Small evangelisation team
  • Tourism and Topography in Britain - Humanities project team
  • Web Office - Small departmental team
  • Wireless Strategy - Small computing service team
  • York Centre for Complex Systems Analysis - Large cross departmental team
  • YorkShare HQ Revamp - Small development team



In some of the cases above, some have barely used the tools, some already have some systems in place whilst others have if anything being using the tools too much, investing lots of time and effort in adding and linking content.

Tuesday, 2 February 2010

The Tools

List of tools to be used