Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

04 April 2013

More Maths for MOOCs!

The guys behind connectivist MOOCs seem to be against the teaching atomic, well-defined concepts, which is all well and good, but I think those same concepts might inform their theories a bit better.

This morning, I received a "welcome to week 4" message from the organiser of the OU MOOC Open Education.  He comes across as a really nice guy on both email and video, which makes it a lot more difficult to criticise, but one thing he said really caught my attention as indicative on the problems with the "informal education" model that the connectivist ideologues* profess. (* I refuse to call them theorists until they provide a more substantial scientific backing for their standpoint -- until then, it's just ideology.)

So, the quote:
"As always try to find time to connect with others. One way of doing this I've found useful is to set aside a small amount of time (15 minutes say) and just read 3 random posts from the course blog h817open.net and leave comments on the original posts."
Don't get me wrong, I commend him on this.  Many MOOC organisers take a step back and stay well clear of student contributions, for fear of getting caught up in a time sink.  No, my problem is the word "random" coupled with the number "3".

The cult of random

There is a scientific truism about random: when the numbers are big enough, random stuff acts predictably.  You can predict the buying patterns of a social group of thousands well enough to say that "of this 10000 people, 8000 will have a car", or the like.

The best examples, though, come in the realm of birthdays.

If you take the birthdays (excluding year) of the population of a country (eg the UK) and plot a graph, you'll get a smooth curve peaking in the summer months and reaching its lowest in the winter months.  Now if you take a single city in that country (eg London), you'll find a curve that is of almost indistinguishable shape, just with different numbers.  Take a single district of the city, and the curve will be a similar shape, but it will start to get "noisy" (your line will be jagged).  Decrease to a single street (make it a large one) and the pattern will be barely recognisable, although you'll probably spot it because you've just been looking at the curve on a similar scale.  Now zoom down to the level of a single house... the pattern is gone, because there aren't enough people.

In physics, this is the difference between life on the quantum scale and the macro scale.  Everything we touch and see is a collection of tiny units of matter or energy, and each of those units acts as an independent, unpredictable agent, but there are so many of these units that they appear to us to function as a continuous scale, describing a probability distribution like in the example of birthdays. Why should I care whether any individual photon hits my eye if the faintest visible star in the night sky delivers 1700 photons per second.  The computer screen I'm staring at now is probably bombarding me with billions as we speak. The individual is irrelevant.

But massive means massive numbers, right?

I know what you're thinking -- with MOOC participants typically numbering in the thousands, these macro-scale probabitimajigs should probably cut in and start giving us predictability.  Well yes, they do, but not in the way you might expect, because MOOCs deal with both big numbers and small numbers.

Again, let's look at birthdays.

Imagine we've got 365 people in a room, and for convenience we'll pretend leap years don't exist (and also imagine that there aren't any seasonal variations in births and deaths).

What is the average number of people born on any given day?
Easy: one.

And what is the probability that there's someone born on every day of the year?
This one isn't immediately obvious, but if you know your stats, it's easy to figure out.

First we select one person at random.
He has a birthday that we haven't seen yet, so he gets a probability of 1, whatever his birthday is OK.
Now we have 364 people, and 364 target days out of 365.
Select person 2 -- the chance he has a birthday we haven't seen yet is 364/365.
Person 3's chance of having a birthday we haven't seen is 363/365... still high.
...
but person 363's chance is 3/365, person 364's chance is 2/365 and person 365's chance is 1/365.

To get the final probability of 1-person-per-day-of-the-year, we need to multiply these:
1 x 364/365 x 363/365 x ... x 3/365 x 2/365 x 1/365

Mathematically, thats
365! / (365^365)
or
364! / (365^364)

It's so astronomically tiny that OpenCalc refuses to calculate the answer, and the Windows Calculator app tells me it's  "1.455 e-157" -- for those of you who don't know scientific notation, that "e minus" number is the number of zeros, so fully expanded, that would be:
0.000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 000 145 5

...unless I've miscounted my zeros, but you can see clearly that the chances of actually getting everyone with different birthdays is pretty close to zero.  It ain't gonna happen.

A better statistician than me would be able to predict with some accuracy the number of dates with multiple birthdays, the number of dates without birthdays etc, but not which ones they would be.

The reason we have these gaps is that we two relatively high numbers and one relatively low number.  That gives us a predicatable probability distribution with predictable gaps.

Now, Martin Weller tries to look at and comment on 3 random blog posts when he connects.  Now let's imagine that this was a policy that everyone adhered to without exception (either above or below).

OK, so let's imagine our MOOC has 365 students (so that we can continue using the same numbers as above) and that for each post we put up we comment on 3.  What's the chances that every blog post gets at least one comment?

Well this is pushing my memory a bit, cos I haven't done any serious stats since 1998, and it's all this n-choose-r stuff.  Oooh.... In fact, I can't remember how to do it, but it's still going to be a very small probability indeed.

In order to get a reasonable chance of everybody getting at least one comment, you need to get rid of small numbers entirely, and require high volumes of feedback for each user.  But even though we're talking unfeasibly high volumes here, it still doesn't guarantee anything.

Random doesn't work!

You cannot, therefore, organise a "course" relying entirely on informal networking and probability effects -- there must be some active guidance.  This is why Coursera (founded by experts in this sort of applied statistics) uses a more formalised type of peer interaction.

The Coursera model

Coursera's model is to make peer interaction into peer review, where students have to grade and comment on 5 classmates' work in order to get a grade for their own work.  The computer doesn't do this completely randomly, though, and assigns review tasks with a view to getting 5 reviews for each and every assignment submitted.

Now in theory, their approach is perfect, because even if they don't achieve a 100% completion rate/0% dropout rate, you should be able to get something back.  However, their execution was flawed in a way which convinces me that Andrew Ng didn't write the algorithm!

You see, the peer reviews appear essentially to be dealt out like a hand of cards -- when the submission deadline is reached, all submissions are assigned out immediately.  Each submitter gets dealt 5 different assignments, each assignment is dealt out to 5 different reviewers.  It doesn't matter when you log in to do the review -- the distribution of assignments is already a fait accompli.

How did I come to this conclusion?  Well, after the very first assignment in the Berklee songwriting course, I saw a post on the course forums from someone who had received no feedback on his assignment.  I immediately checked mine: a full set of 5 responses.

Even though they had distributed the assignments for peer review in a less random fashion, they did nothing to account for the dropout rate, even though the dropout rate is reportedly predictably similar across all MOOCs -- and not only similar, but very high in the first week or two.  So statistically speaking, gaps were entirely predictable.

What Coursera should have done....

The problem was this "dealing out" in advance.  If they'd done the assignment redistribution on a just-in-time basis.  When a reviewer starts a review, the system should assign one with the minimum number of reviews.  No-one should receive a second review until everybody has received one, and I definitely shouldn't have received 5 when some others were reporting only getting 3, 2 or even none.

"Average" is always a bad measure

As a society, we've got blasé about our statistics, and we often fall back on averages, and by that I mean "mean".  But if the average number of reviews per assignment is 4.8 of a targeted 5, that's still not success and it's still not if it doesn't mean that everyone got either 4 or 5.

Informal course organisation will never work, no matter how "massive" the scale, because the laws of statistics don't work like people expect them to -- they don't guarantee that everyone gets something -- they guarantee that someone gets nothing.

29 May 2012

Authentic: long v short pt 5

Today was a wee bit frustrating.  I spent a solid chunk of time trying to get the ntlk data module installed, and with it the file english.pickle that would have allowed me to do part-of-speech (POS) tagging.  This would have made it almost trivially easy to eliminate the proper nouns and get a genuine look at the real "words" that are of interest to the learner.

Ah well, it looks like it wasn't meant to be.

So I started working towards custom code to eliminate the proper nouns manually, something which would be handy in the future anyway.  The first step was to identify some candidates for further inspection, and seeing as I'm working with English, that's pretty easy: if it's not all in lower case, there's something funny about it.  I wrote the code to identify all the tokens (words) that contained capitals.  Yes, at this point I could have checked whether it was the start of the sentence or not, but that wouldn't have really helped, because proper nouns occur at the start of sentences too, so i'd still need to check.

When I generated my set of candidates, though, it was a little long.  For The 39 Steps, I was looking at 919 tokens to check manually, and that's a fairly short book.  As I'm doing this for fun, it seemed like checking that many would be a little bit boring, particularly in longer books.  (I later checked the candidate set for the 3 books in total, and it turned out to be over 3000 words, which is more than my time's worth.)

My first quick test then was to have a look at the difference in figures.  Eliminating every single item with any capitals in it drops the type:token ratio in The 39 Steps from 14.48% to 13.14% -- that's almost a a 10% drop (it's 1.35 percentage points, but it's 9.27 percent).  Before properly addressing the proper nouns, I wanted to see how big a difference this crude adjustment makes to the figures.  It seemed just a little too high to realistically be led by proper nouns alone.  But can that be?  I mean, how many words are likely to occur only at the start of sentences?

So on I went, hoping that the data I could generate at this stage would start to shed some light on this figure.

The first graph I produced showed me the running type:token ratios and introduction rates for both the full token set, and the token set with non-lowercase words eliminated:
The two pairs of lines follow each other pretty closely, getting closer together as they progress.  But in order to start getting a clear idea of what was going last time, I had to go to another level of abstraction and measure some useful differences.  So here is the difference between the running ratios for all words and lowercase only, and the corresponding difference in introduction rates:
Now you'd be forgiven for thinking that the difference is diminishing here -- I was fooled into thinking the same thing, but then I realised I was dealing with numbers here rather than proper stats, and I redid the analysis but with a difference in percentage:
The overall running type:token ratio does indeed decrease, but it halves (20% down to 10%) then stabilises.  The introduction rate, on the other hand, is all over the place -- there's no identifiable trend at all.  Even subsampling my data didn't give any clear and understandable trends (and since I'm using a desktop office package for my analysis it's a bit of faff to do the resampling automatically -- it's just further proof that I need to get myself familiar with the statistical analysis tools for Python (eg numpy), but my head's full with the NLTK stuff for now, so I'll leave the improved statistical stuff for
another time).  Here's the same graphs, but with 2000 word samples instead of 500 word samples:

So not promising, really.  Still no stable, identifiable trends.

Books as a series
But I had all the infrastructure in place now, so I figured I might as well rerun the analysis on the 3 books as a single body and see what came out.  Let's just go straight to the relative difference between the lines for all words and eliminating all words not entirely in lower case:
Oooh... now where did I leave those figures on where the individual books started...?  44625 and 152034, and there's a notable period of high difference (20-30%) from about 45000 words, and that massive spike you seem on the graph -- which is actually a 63.64% difference -- occurs from 152000-152500.

Bingo: we've got decent support for Thrissel's suggestion that a lot of proper nouns are introduced early on in... at least some novels.

Not the sort of information I was originally looking for, but actually quite interesting.  It's kind of turning the project in a slightly different direction than I had planned.  I'll just have to go with the flow.

What I did wrong today
One of the minor irritations of the day was when I started writing up my results, and after having done the coding, data generation and analysis, I realised a fairly simple refinement I could have made.  It was a real *palmface* moment: I could have simply taken my first list of candidate proper nouns and eliminated any candidates that also appeared completely in lower case.  Having done that, I would have been left with a much shorter list of candidates, and it may well have been worth my time manually checking the results.

>sigh<

But of course, that's as much the point of the exercise as anything: to work through the process and the problems and to start thinking about what can be done better.

It also occurs to me now that I also managed to eliminate every single occurrence of the word I from the books!  Quite a fundamental error, even if it only made a minute difference to the final ratios.

Perhaps I'm being a little too "hacky" in all this.  I'll have to pick up my game a bit soon....

28 May 2012

Authentics: long v short - pt 4

Well, I had a nice weekend and visited some friends in Edinburgh for a wedding.  The weather's too good to spend too much time inside, so I'll just write up a few more tests then go and enjoy the sunshine.

Multiple books in a series
Today's figures come from two of the books I've already mentioned -- John Buchan's The 39 Steps and Greenmantle, and the next book in the series: Mr Standfast.

Again, the graphs at different sample sizes show different parts of the dataset more clearly than others:
The graph as 1000 words is too unstable to clearly identify the end of the the first book, and the start of the second book is only identifiable because the introduction rate is greater than the running ratio for the first time in any of my tests.

The 2500 and 5000 word samples give us a clear spike for the end of the first book, but the end of the second book is obscured slightly by noise, and becomes clearer again in the 10000 and 25000 word sample sizes, although the end of the first book is completely lost by the time we reach the 25000 word sample graph.

Having done all that, I went back and verified the peaks matched the word counts -- The 39 Steps is 44625 words long, and Greenmantle is 107409 words long, so ends with the 152034th word.  The peaks on the graphs all occur shorlty after 45000 and 150000.

It was the first graph, from the 1000 word sample set, that piqued my curiosity.  Having spotted the two lines intersecting at the start of the third book, I decided to check the difference between the running type:token ratio and the introduction rate, and I graphed that.  At 1000 word samples, there was still too much noise:
However, given that I already knew what I was looking for, I could tell that it showed useful trends, and even just moving up to 2500 word samples made the trends pretty clear:
Going forward, I need to compare the difference between books in a series, books by the same author (but not in a series) and unrelated books, and I believe that the difference between the running type:token ratio and the introduction rate may be the best metric to use in comparing the three classes.

Problems with today's results
I can't rule out that the spikes in new language at the start of each book aren't heavily influenced by the volume of proper nouns, as Thrissel suggested, so I'm probably going to have to make an attempt at finding a quick way of identifying them.  The best way of doing this would be to write a script that identifies all words with initial caps in the original stream, then asks me if these are proper nouns or not.

By treating the three books as one continuous text in the analysis, it looks like I've inadvertently smoothed out the spike somewhat at the start of each book.  In future I should make sure individual samples are taken from one book at a time so that the distinction is preserved.

24 May 2012

Authentic materials for learners: long form or short form? (pt 2)

Well, as I was saying in part I, I've always claimed novels are easier for the learner than short stories, and I was wanting to back up my claims with some figures.  So for my initial investigation I fired up my copy of the free TextSTAT package and away I went.

I was talking about the type:token ratio last time, and that seemed as good a place as any to start.  I managed to skip the logical first step, which would have been to compare a novel and a short story, but I'll have to come back to that later.

What I started with was an Italian novel, but I got figures that were too high to be useful.  One of the problem with languages such as Italian is that they write some of their clitics in the same written word as the main word (EG "to know (someone)" -> conoscere; "to know me" -> conoscermi), increasing the type:token ratio significantly.  You've also got the problem that it has verb conjugations and it drops subject pronouns in most situations.  Overall Italian (and Spanish and Catalan, among others) would be a bad choice for a demonstration language.  Today, I'm using English as it's a very isolating language -- the only common inflections are past-tense-ed, second-person-present-s and plural-s.  This makes it easy to get a reasonably accurate measure of the lexical variety without any clever parsing.  I will most likely use French at some point too, because while it is not as straightforward as English in that sense (it's got a lot of verb conjugation going on), it doesn't have the same clitics problem as Italian, and the French don't drop their pronouns.
Today's findings: 1 - running ratios
I decided to look at how the type:token ratio changes as a text proceeds.  I wanted to measure this chapter by chapter, counting the types and tokens in chapter 1, then loading the second chapter into the concordance and checking the type:token ratio for chapters 1 & 2 combined, then 1, 2 & 3 etc.  I realised, however that it would be more efficient to load all chapters into memory at the same time and work down from the other end: all chapters, then close the last chapter and take the figures again, then close the second last chapter and take the figures again.
In the end, I got a nice little graph (using LibreOffice) that showed a marked tendency to decreasing type:token ratio as the books progressed:
The x-axis shows the chapter number, the y-axis shows the type:token ratio (remember, this is the type:token ratio for the entire book up to and including the numbered chapter).  Notice how the type:token ratio halves by around the 6th or 7th chapter.

So by one measure, the longer the novel is, the easier it would appear to be.
Today's findings: 2 - introduction rates
I figured I could go a bit deeper into this without generating any new data.  What I wanted to look at now was how much new material was introduced in each chapter -- ie. a ratio of new types to tokens. It's easy enough to do -- I could obtain the number of new types in any given chapter by deleting the running total at the previous chapter from the running total at the current chapter.
The graph I got was even more interesting than the last:
While the running ratio halves after 6 or 7 chapters, the introduction rate halves after only 2-4!  It certainly looks like each chapter will on average be easier than the last.
One curious feature is the large uptick at the end of the children's novel Laddie (green).  This illustrates one quirk that the learner should always bear in mind: kids books are often actually more complicated linguistically than adults' books, as the author on some level seeks to educate or improve the person reading.  The author of this book seems to have kept the language consistently simple through most of the book, but realising he was coming to the end, crammed in as much complexity as possible.

Another curious feature is that the figures claim no new vocabulary is introduced in the fourth chapter of The 39 Steps (yellow).  While this is theoretically possible, its more likely that it's ...ahem... experimenter error, which a quick look at the actual figures verifies: chapters 3 and 4 are listed in my output as being exactly the same length, which is more than a little unlikely.  It looks like I loaded the same chapter twice...
Further analysis
Notice that in both graphs, the figures are the same at chapter one.  This is to be expected, as every type encountered in the first chapter is encountered for the first time in the book (by definition).

So what happens if we stick the running ratio of type:token against the introduction rate of new types?

This:
So while the overall type:token ratio continues to fall notably from the 10th to the 20th chapter, suggesting decreasing difficulty, the introduction rate gets fairly erratic by around the 10th chapter (despite still tending downwards), so perhaps there is a limit after which it is not safe to assume that each chapter is a difficult as the last.

Perhaps the measure of efficiency is related to the difference between the running ratio and the introduction rate, and once that gap starts to narrow, there is no advantage?

Problems with today's findings
This was a first exploratory experiment, so I didn't conduct it with a whole lot of rigour.  Here are the main factors affecting todays results:
  1. I didn't eliminate common words -- it is impossible to see from the figures I have how many of the types introduced at any stages are ones we would expect learners to know already and how many will be genuinely new to them.
  2. When examining Pride and Prejudice and The 39 Steps, I hadn't told the concordancer to ignore case, so anything appearing at the start of a sentence and in the middle would be counted as two types -- eg that and That.  (It was the first time I'd used TextSTAT and I hadn't realised it defaulted to case-sensitive -- I won't make that mistake again.)
  3. The length of chapters varies significantly from book to book and even from chapter to chapter within books, so the lines are not to scale with each other, and each individual line is not in a continuous scale with itself.  The graphs, though presented in a line, are arguably not true line graphs, as they occur from samples arbitrarily dispersed.
Accounting for these problems in the future
  1. There are plenty of frequency lists on the net, so I'll be able to eliminate common words without any real difficulty.
  2. The case sensitivity issue, now that I'm aware of it, will not be a problem.
  3. When I ran the initial data, I was using TextSTAT as my installation of Python and NLTK was playing up (I had too many different versions of Python installed, and some of the shared libraries were conflicting).  I've now got Python to load NLTK without problems, so I can do almost any query I want.  Future queries will be sampled regularly after a specific number of words.
Experiments to carry out
At some point I'm going to want to go back and compare short stories with novels, but for now I'm going to head a little further down the path I'm on.

My first task is to work out a decent sampling interval: ever 1000 words? 5000? 10,000? 50,000?  I'll run a few trials and see what my gut reaction is -- that should be the next post.  (It might even prove that the chapter is the logical division anyway -- after all, it divides subjects, which would indicate different semantic domains...)
I also want to look at what happens when we look at sequels after each other.  Those of you familiar with John Buchan will notice that I've included such a pair as individual novels here -- The 39 Steps and Greenmantle.  I might include initial findings from this next time, as they'll determine my next step.
After this I'll either move on to looking at more pairs of original book + sequel (to look for a generalisable pattern), looking at longer serieses of books (to see if they get continually easier) or comparing book-and-sequel to two different books from the same author (to see if any perceived benefits from reading a book and its sequel are just coincidence and really only because of the author).

Caveat emptor
Remember, though, that this little study is never going to be scientifically rigorous, as I don't really currently have the time to deal with the volume of data required to make it truly representative.  However, it's nice to think how big a job this would have been before computers made this sort of research accessible to the hobbyist.  Many thanks to the guys who wrote the various tools I'm using -- your work is genuinely appreciated.