Brains of different people are not in a one to one correspondence, they do not have the same number of cells and even if they had, it is not known if the same information will get encoded in the exact same cell.
This is true, of course, but irrelevant to this work. At the level they're working at (fMRI scans, which have a resolution on the order of 0.5-4mm or so depending on the temporal resolution, etc.), you can't resolve individual cells anyways so you don't have to be concerned about those kinds of individual variations.
Visual activity in many parts of the brain follows a retinotopic map, where activity in nearby locations on the retina are processed in nearby regions of the brain. So, while you would have to calibrate some details, a lot of things would be constant between brains.
Inter-subject variability is a huge problem in this kind of work.
As you say, they're operating at a much larger scale than individual cells (100k-1m cells in each voxel). Likewise, some early visual processing areas are broadly organized kind of like big, noisy bitmaps on the surface of the brain.
But for sophisticated machine learning-style analyses like these, the gross differences in representation and morphology (especially at higher processing levels in the brain) make it very hard to pool the data across multiple people. That's why they're preferring to use many sessions from a small number of participants rather than a single session from many participants (the standard approach).
[I worked on applying machine learning methods to fMRI for my PhD]
Indeed and this is more true for the retinotopic map. For cognitive there seems to be quite a bit of variability. In fact its not a trivial task to distinguish a cognitively active brain from one that is in a waking but resting state.
The comment that you quote wasn't made with reference to their work specifically but to any sufficiently accurate technology that can read through your eyes via the brain.
This is true, of course, but irrelevant to this work. At the level they're working at (fMRI scans, which have a resolution on the order of 0.5-4mm or so depending on the temporal resolution, etc.), you can't resolve individual cells anyways so you don't have to be concerned about those kinds of individual variations.
Visual activity in many parts of the brain follows a retinotopic map, where activity in nearby locations on the retina are processed in nearby regions of the brain. So, while you would have to calibrate some details, a lot of things would be constant between brains.
http://en.wikipedia.org/wiki/Retinotopy