The WSJ has an article today called "Insurers Test Data Profiles to Identify Risky Clients".
It turns out that "risky behavior" can be inferred from scanning the internet (including social web sites) for "bread crumbs" (or maybe cookie crumbs) you leave about yourself. What sort of bread crumbs are these you might ask?
For example, hunting permits, boat registrations and property transfers, purchase histories, credit data, forum discussions, blogs, an so forth - all public information or information you allowed someone to collect about you. Remarkably, according to the article, this data can be used to tell things like if you like gourmet food, if you exercise regularly, or sit on the couch, commute long distances, watch too much TV, and many other things. There is a lengthy discussion of how this works in the article.
You'll probably also be surprised to know that is not a new idea, either.
Twenty years ago companies like auto insurers, according to the link, began to use things like credit score to decide how to price your policy (the higher the credit score the less likely you are to file claim).
Now the really interesting part of all this is that life insurers are thinking about how this can replace things like a "blood test" to determine what kind of "risk" you are in the case of insurance. I think this replacement of an actual "blood test" with an internet "risk" test is itself risky business on the part of the insurance companies.
So let's talk about risk as it relates to this sort of data mining. A few days ago I wrote this article about "Cholesterol, Heart Disease and Magical Thinking". People do not understand how people like epidemiologists calculate risk - and that's a big problem. You hear about risk all the time: on TV, in ads, from friends, from doctors. Don't do this or that because its "risky".
First of all, what is risk? Well, for one thing its not a predictor that something will happen. A predictor is something that, when we observe it, tells us the a high degree of certainty that some other corresponding event will occur. For example, a clap of thunder can be predicted from the observation of a lightning bolt. Risk is also not a cause of something. Causality is represented by a direct link between two events, i.e., lightning and thunder. We can say the lightning bolt caused the clap of thunder to occur.
But risk is different. So how is this kind of risk defined? What does it mean?
Well, in epidemiology risk factors are calculated as follows:
We take a statistically significant group of people (you can use common sense here - for something like heart disease you wouldn't study just five people - you'd study a large number). Just how large a number is not really important here, all we need to know is the number is large enough for statistical purposes.
We'll pretend in this post that 100 people are subjects in the study because math with 100 is relatively easy.
So let's say (and we're making this up) that 20 people out of our 100 subjects have had heart attacks. That's 20 / 100 = .20 = 20%. So we say that in general you have a 20% risk of heart attack.
Let's also say that 25 people in our example buy pants with a waste size of 40 or above and we'll pretend that 15 people in this "large pant size group" also have had heart attacks.
So the number of people that have a "large pants size" and have had a heart attack is 15, or 15 / 100 or .15 or 15% of the population.
If we divide the 15% (people who purchase "large pants" and have had an heart attack) by the 20% that just have had a heart attack we get .75 or 75% risk factor that if I buy large pants I will have had a heart attack.
But what does this risk factor really mean?
Nothing concrete. It does not tell anyone what you will do - it just says that when a lot of people get together there is a chance that something will occur. A risk factor represents this numerical chance that something might happen based on examination of a large group. (Chance here is a number between zero and one, commonly shown as a percentage, i.e., .1 = 10%.) Sort of like saying 10% of the people at a baseball game buy hot dogs. We don't know which people will buy hot dogs but we can generally assume that for any given baseball game about 10% will buy hot dogs - everything else being equal (for example, there are no sales of hamburgers that day). This is why stadium vendors can buy just about the right amount of food so none is wasted and they don't run out.
So using purchase histories, information about permits, and so on statisticians can develop an entire profile about you that tells them what sort of risk you are relative to whatever insurance policy you are applying for. But, if you are clever, you will also realize something else.
Just because you purchase large pants doesn't mean you wear them.
For example, you may have an elderly relative at home who you care for and you go online to purchase their clothing for them and not yourself.
This is a big difference here between a true epidemiological study and "skimming and mining" data from internet sites and data providers. In a true epidemiological study we can have definite links between buying and wearing the pants, i.e., we can include in our study the notion of collecting definite data. Here we don't know who is using the products we are buying. We're assuming the purchaser is the user - and we all know what assuming does.
So to some degree I see this entire process as "magical thinking" on the part of those insurance companies correlating internet data with personal risk, i.e., the insurance companies. Correlation means, in this case, that when one thing happens there is an observed relationship with some other thing happening. A correlation is an observation.
Dogs make correlations: If I walk to the container holding the dog food they think I am going to feed them - so they stick close by me. The dog mind predicts that I will feed them when I do this. But walking to the dog food container does not cause me to feed them. Similarly if I walk by the dog food container all the time and don't feed them the dogs will soon realize that their correlation is not useful and abandon it.
So one problem here is that, using this sort of system, some behavior you have, for example purchasing pants for an elderly person you take care of, may be correlated with you instead of the actual user of the purchase, i.e., the elderly relative.
Another potential problem here is that even though there may be invalid individual correlations a system like this is making about you the overall predictive ability of the model may still work. For example, caring for an elderly relative might cause you a lot of stress and you have a heart attack because of it. Effectively this becomes some form discrimination.
Hence the bread crumbs you leave behind may be leading others on a false trail.
I have been involved in high tech, graphic arts, computer software and hardware design for more than 40 years. I've been blogging about vaping since early 2009. I work on advanced robot vision, 3D, SONAR, LIDAR, and software technology. I own my own business. I have set up this blog to talk about who I am, what I do, and to publish my opinions...
Search This Blog
Friday, November 19, 2010
Thursday, November 18, 2010
Copyright Trolls
Several days ago I wrote"Patents, Trash Mobs and Apple Pie" about an editor at a publication called "Cooks Source." The post was about the republishing of articles written and owned by others. This was interesting to me because it illustrated how law and the web interact.
We now come upon a company called Righthaven, LLC.
Righthaven came to light a while back by suing a couple of bloggers in federal court who "reposted" stories from a publication called the Las Vegas-Review Journal (LVRJ). In both cases the reposts contained some or all of the original LVRJ story with credit. One blogger was Mary J. Santilli of Boston and an "American Idol" fanatic posted a full LVRJ story on her blog with credit. Another was Allegra Wong, also of Boston, who wrote a story from her cat's perspective about a fire that killed some birds - again providing credit but only using a portion of the story.
Since then Righthaven has filed additional federal lawsuits in about 100 or so cases. Typically the cases ask for "damages of $75,000 and forfeiture of website names" according to the Las Vegas Sun.
So let's see, 100 x $75,000.00 US = $7,500,000.00 USD.
The suing proceeded until Righthaven came across Reality One Group, Inc. Group One fought back and won based on this decision. The decision reads in part: "The Fair Use doctrine states in pertinent part that “the fair use of a copyrighted work, . . . for purposes such as criticism, comment, [or] news reporting . . . is not an infringement of copyright.” 17 U.S.C. § 107. In determining if an alleged infringement is a fair use of the copyright, district courts consider several factors including: (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work. 17 U.S.C. § 107; see also A&M Records, Inc. v. Napster, Inc., 239 F.3d 1004 (9th Cir. 2001). "
The judge goes on to discuss this indicating that while Group One may be a commercial entity using LVRJ's content it meets the "Fair Use" exception to copyright law because A) the use is factual new reporting and commentary, B) Group One use only eight of thirty sentences in an article (about 26%), and C) the use does not dilute the original copyright holders (LVRJ) market. For me this decision is dead on with standard "Fair Use" doctrine (link here).
Another judge, Robert Johnston, has questioned Righthaven's court costs related to these suits. Righthaven, for example, demands costs in these suits and its in-house counsel charges $160 - $190 USD per hour.
All this said we have to now step back somewhat and view this from a different perspective.
While its clear that posting an entire article from a publication like LVRJ would be infringement what's not so clear is the legality of the process that LVRJ has constructed with Righthaven to attack supposed infringement.
Under the Digital Millennium Copyright Act (DMCA) a blog host, such as blogspot in the case of this blog, should register a "takedown agent". A "takedown agent" is a person who is to be notified if infringing material appears on the site. If you do not have such a person registered the actions of an entity like Righthaven can be more problematic. Under the DMCA if a blog has a registered takedown agent Righthaven must notify the agent of any infringement. The agent must then take steps to "take down" the infringing material.
In the case of Google, typically this means simply removing the material, though there are a lot of complex legal elements related to this (see this). A more complete discussion of your rights relative to "online-freedom" can be found here.
In defense of another Righthaven target the Electronic Freedom Foundation (EFF) has filed a counter suit against Righthaven alleging that it is a "Copyright Troll" that seeks to extract “... windfall recoveries of statutory damages and to exact nuisance settlements” from its targets.
So what's the bottom line to all this?
Clearly in a case like "Cooks Source" copying an entire article into your blog or publication without permission is wrong.
Its also clear that trolling for cases like this with the intended purpose of extracting "nuisance settlements" is also wrong. (It will be interesting to follow this counter claim - my guess is that the US legal system will ultimately reject the notion of nuisance claims for a number of reasons).
The real losers here are you and I (unless you're a lawyer).
The creation and filing of nuisance lawsuits is an all too common practice of which I myself have been a victim over the years. What these articles don't say is that any defense, no matter how simple, is likely to cost tens of thousands of dollars in addition to whatever claim is made against you (unless the EFF comes to your rescue).
I think the US legal system includes the implicit presumption that the filer of claims like Righthaven's have substantial merit - they wouldn't have sued if they didn't, right? This bias comes from a history of law developed over several hundred years before the electronic age. You and I did not file claims like this except in exceptional circumstances - and most publishers (like newspapers) respected copyright and the law - lawsuits were the realm of big business who could afford them.
Today, however, the mere threat of a lawsuit is substantial - far beyond the resources of most individuals or small businesses.
And this argument goes to both sides: In the case of Cooks Source I mentioned at the beginning of the article the original author could have filed a federal lawsuit against Cooks Source claiming infringement. However, this would also have cost tens of thousands of dollars - regardless of any success or failure.
The "takedown" portion of the DMCA is a step in the right direction - but it doesn't go far enough.
There needs to be a simple, straightforward legal mechanism to handle the first level of these types of claims without need of lawsuits, lawyers and judges. This would allow you and I to handle issues that arise without the involvement of legal trolls looking to benefit from the hard work of others.
The bottom line is this:
In the case of Cooks Source using the "Tale of Two Tarts" article without the permission of Monica Gaudio wouldn't it be better to have this resolved through a simple legal means, e.g., arbitration or the like, rather than have a legal trolls collecingt tens of thousands of dollars from both sides?
Its wrong to steal copyrighted materials - but I think its just as wrong to profit beyond the original value and scope of the material.
My guess is that the "Tale of Two Tarts" did not generate nearly enough money for Monica Gaudio to pay for a federal lawsuit against Cooks Source.
Damages and costs should be limited to the real value of the materials and harm in question.
We now come upon a company called Righthaven, LLC.
Righthaven came to light a while back by suing a couple of bloggers in federal court who "reposted" stories from a publication called the Las Vegas-Review Journal (LVRJ). In both cases the reposts contained some or all of the original LVRJ story with credit. One blogger was Mary J. Santilli of Boston and an "American Idol" fanatic posted a full LVRJ story on her blog with credit. Another was Allegra Wong, also of Boston, who wrote a story from her cat's perspective about a fire that killed some birds - again providing credit but only using a portion of the story.
Since then Righthaven has filed additional federal lawsuits in about 100 or so cases. Typically the cases ask for "damages of $75,000 and forfeiture of website names" according to the Las Vegas Sun.
So let's see, 100 x $75,000.00 US = $7,500,000.00 USD.
The suing proceeded until Righthaven came across Reality One Group, Inc. Group One fought back and won based on this decision. The decision reads in part: "The Fair Use doctrine states in pertinent part that “the fair use of a copyrighted work, . . . for purposes such as criticism, comment, [or] news reporting . . . is not an infringement of copyright.” 17 U.S.C. § 107. In determining if an alleged infringement is a fair use of the copyright, district courts consider several factors including: (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work. 17 U.S.C. § 107; see also A&M Records, Inc. v. Napster, Inc., 239 F.3d 1004 (9th Cir. 2001). "
The judge goes on to discuss this indicating that while Group One may be a commercial entity using LVRJ's content it meets the "Fair Use" exception to copyright law because A) the use is factual new reporting and commentary, B) Group One use only eight of thirty sentences in an article (about 26%), and C) the use does not dilute the original copyright holders (LVRJ) market. For me this decision is dead on with standard "Fair Use" doctrine (link here).
Another judge, Robert Johnston, has questioned Righthaven's court costs related to these suits. Righthaven, for example, demands costs in these suits and its in-house counsel charges $160 - $190 USD per hour.
All this said we have to now step back somewhat and view this from a different perspective.
While its clear that posting an entire article from a publication like LVRJ would be infringement what's not so clear is the legality of the process that LVRJ has constructed with Righthaven to attack supposed infringement.
Under the Digital Millennium Copyright Act (DMCA) a blog host, such as blogspot in the case of this blog, should register a "takedown agent". A "takedown agent" is a person who is to be notified if infringing material appears on the site. If you do not have such a person registered the actions of an entity like Righthaven can be more problematic. Under the DMCA if a blog has a registered takedown agent Righthaven must notify the agent of any infringement. The agent must then take steps to "take down" the infringing material.
In the case of Google, typically this means simply removing the material, though there are a lot of complex legal elements related to this (see this). A more complete discussion of your rights relative to "online-freedom" can be found here.
In defense of another Righthaven target the Electronic Freedom Foundation (EFF) has filed a counter suit against Righthaven alleging that it is a "Copyright Troll" that seeks to extract “... windfall recoveries of statutory damages and to exact nuisance settlements” from its targets.
So what's the bottom line to all this?
Clearly in a case like "Cooks Source" copying an entire article into your blog or publication without permission is wrong.
Its also clear that trolling for cases like this with the intended purpose of extracting "nuisance settlements" is also wrong. (It will be interesting to follow this counter claim - my guess is that the US legal system will ultimately reject the notion of nuisance claims for a number of reasons).
The real losers here are you and I (unless you're a lawyer).
The creation and filing of nuisance lawsuits is an all too common practice of which I myself have been a victim over the years. What these articles don't say is that any defense, no matter how simple, is likely to cost tens of thousands of dollars in addition to whatever claim is made against you (unless the EFF comes to your rescue).
I think the US legal system includes the implicit presumption that the filer of claims like Righthaven's have substantial merit - they wouldn't have sued if they didn't, right? This bias comes from a history of law developed over several hundred years before the electronic age. You and I did not file claims like this except in exceptional circumstances - and most publishers (like newspapers) respected copyright and the law - lawsuits were the realm of big business who could afford them.
Today, however, the mere threat of a lawsuit is substantial - far beyond the resources of most individuals or small businesses.
And this argument goes to both sides: In the case of Cooks Source I mentioned at the beginning of the article the original author could have filed a federal lawsuit against Cooks Source claiming infringement. However, this would also have cost tens of thousands of dollars - regardless of any success or failure.
The "takedown" portion of the DMCA is a step in the right direction - but it doesn't go far enough.
There needs to be a simple, straightforward legal mechanism to handle the first level of these types of claims without need of lawsuits, lawyers and judges. This would allow you and I to handle issues that arise without the involvement of legal trolls looking to benefit from the hard work of others.
The bottom line is this:
In the case of Cooks Source using the "Tale of Two Tarts" article without the permission of Monica Gaudio wouldn't it be better to have this resolved through a simple legal means, e.g., arbitration or the like, rather than have a legal trolls collecingt tens of thousands of dollars from both sides?
Its wrong to steal copyrighted materials - but I think its just as wrong to profit beyond the original value and scope of the material.
My guess is that the "Tale of Two Tarts" did not generate nearly enough money for Monica Gaudio to pay for a federal lawsuit against Cooks Source.
Damages and costs should be limited to the real value of the materials and harm in question.
Wednesday, November 17, 2010
GigaPans and Xeikons
I have always been interested in GigaPan.
GigaPan is a system that allows you to create an enormously detailed image by stitching together a large number of high resolution images - each of just a tiny fraction of the whole picture. Special camera attachments and software tools allow you to take these pictures and you can also use software to stitch together images you have lying around. There is a web viewer that allows you to zoom around in the stitched result as if it were one giant picture.
The viewer works a lot like the Google Map viewer. You can zoom in and out from an interplanetary view down to your mailbox.
I found this site at National Geographic - it has some very cool examples (check out the "Pill Bug" if you're not squeamish).
GigaPan is a partnership between Google, CMU, NASA and a few others. The site claims its an extension of the Google Connection Project (site here, but it loads very slowly) whose purpose is "... develop[s] software tools and technologies to increase the power of images to connect, inform, and inspire people to become engaged and responsible global citizens."
The technology is amazing but I am surprised that there aren't a lot of commercial applications for it yet.
One that would seem obvious is "digital pathology" - taking the slides pathologists make from biopsy's and so forth and converting them into GigaPan images. There is a company in Pittsburgh call Omnyx which is developing such a platform - but as far as I can see it does not use GigaPan.
I thought about this a bit and it seems reasonable that a commercial venture would want to make sure that there was sufficient bandwidth to load the images quickly and smoothly - something you could not necessarily guarantee on an regular internet connection.
In a lot of ways this is similar to the Xeikon Digital Press technology that allows images invoked by PPML to be streamed to the press on demand. For the Xeikon press you RIP various elements of the job onto a server available to the press via a network. As the press runs the PPML driving the job calls in assets. The assets are then pulled in over the network by the press.
In the case of the Xeikon there is a much greater demand on performance and reliability of image delivery because the moving paper really requires that the images arrive on time - if they don't the press really has no choice but to stop with an error.
I recall talking to the Omnyx people about this but they seemed very interested in re-inventing the wheel.
I would imagine that another issue for Omnyx is depth of field. Though a GigaPan has tremendous resolution it only has it at a particular focus distance. But a pathologist would probably like to have focus at various distances so that he could see what's effectively "behind" or "in front of" some element on the image.
I bet it would be easy to create a GigaPan viewer that supports a depth of field adjustment allowing you to make an on-the-fly adjustment while you are viewing.
I also noticed on a lot of GigaPans that the focus at high resolution is relatively poor. For example, you have a beautiful mountain scene from a great distance and you can zoom into the specific trees on one part of one mountain. But the focus on those trees is not sharp.
GigaPan has just released a camera mount system that automatically takes a sequence of images (it costs $895.00 US). If you had two and they were synced you could do 3-D. A friend of mine bought one of these - it seems to work quite well.
GigaPan is a system that allows you to create an enormously detailed image by stitching together a large number of high resolution images - each of just a tiny fraction of the whole picture. Special camera attachments and software tools allow you to take these pictures and you can also use software to stitch together images you have lying around. There is a web viewer that allows you to zoom around in the stitched result as if it were one giant picture.
The viewer works a lot like the Google Map viewer. You can zoom in and out from an interplanetary view down to your mailbox.
I found this site at National Geographic - it has some very cool examples (check out the "Pill Bug" if you're not squeamish).
GigaPan is a partnership between Google, CMU, NASA and a few others. The site claims its an extension of the Google Connection Project (site here, but it loads very slowly) whose purpose is "... develop[s] software tools and technologies to increase the power of images to connect, inform, and inspire people to become engaged and responsible global citizens."
The technology is amazing but I am surprised that there aren't a lot of commercial applications for it yet.
One that would seem obvious is "digital pathology" - taking the slides pathologists make from biopsy's and so forth and converting them into GigaPan images. There is a company in Pittsburgh call Omnyx which is developing such a platform - but as far as I can see it does not use GigaPan.
I thought about this a bit and it seems reasonable that a commercial venture would want to make sure that there was sufficient bandwidth to load the images quickly and smoothly - something you could not necessarily guarantee on an regular internet connection.
In a lot of ways this is similar to the Xeikon Digital Press technology that allows images invoked by PPML to be streamed to the press on demand. For the Xeikon press you RIP various elements of the job onto a server available to the press via a network. As the press runs the PPML driving the job calls in assets. The assets are then pulled in over the network by the press.
In the case of the Xeikon there is a much greater demand on performance and reliability of image delivery because the moving paper really requires that the images arrive on time - if they don't the press really has no choice but to stop with an error.
I recall talking to the Omnyx people about this but they seemed very interested in re-inventing the wheel.
I would imagine that another issue for Omnyx is depth of field. Though a GigaPan has tremendous resolution it only has it at a particular focus distance. But a pathologist would probably like to have focus at various distances so that he could see what's effectively "behind" or "in front of" some element on the image.
I bet it would be easy to create a GigaPan viewer that supports a depth of field adjustment allowing you to make an on-the-fly adjustment while you are viewing.
I also noticed on a lot of GigaPans that the focus at high resolution is relatively poor. For example, you have a beautiful mountain scene from a great distance and you can zoom into the specific trees on one part of one mountain. But the focus on those trees is not sharp.
GigaPan has just released a camera mount system that automatically takes a sequence of images (it costs $895.00 US). If you had two and they were synced you could do 3-D. A friend of mine bought one of these - it seems to work quite well.
Tuesday, November 16, 2010
Are We There Yet?
Based on Mark's comment from yesterday I updated my knowledge about E-Paper.
The Palo Alto Research Center (Parc), a research center for Xerox, was well known for inventing the Alto, a precursor of most modern computer systems. The Alto, as the computer was called, consisted of an networking card (Ethernet - invented by Robert Metcalfe), a bitmap display, a mouse, a custom processor, and a hard drive (2.5 Mb).
The "Dover Laser Printer" was also invented around this time (I wrote about it here). It was, as far as I know, the first networked printing device (as well as the first laser printer).
The concept of E-paper was invented by Nicholas K. Sheridon at Xerox Parc in the mid 1970's.
Some interesting material is presented here including a lengthy interview with Sheridon.
Sheridon makes an interesting point: "Much has been written about the incredible myopia of Xerox executives of the time, so I won't go into that except to say that there were numerous other opportunities to enormously expand Xerox's business that were similarly fumbled. Xerox had enough money to create an incredible research lab with top-notch people, but Xerox management could not shake off the copier mentality."
Xerox had literally invented the future of computing at Parc by 1980 or so. Everything you use and take for granted in a computing sense was created there. Supposedly Steve Jobs "stole" the idea to create the Lisa - the precursor of the Macintosh - from Parc after a visit.
But Xerox management could not understand what their research team had invented: they only understood copiers. The proof, of course, is that only the Dover was commercialized by Xerox (initially as the Xerox 9700). They eventually tried to commercialize the Alto as the Xerox Star Office - but it was a dismal failure. Even with the commercialization of the Dover as the X9700, however, the networking and so forth was discarded in favor of a mainframe channel adapter (for communicating with main frames) and an 9-track, reel-to-reel tape drive.
The link covers Gyricon and E-Ink in some detail which were various spin-offs from Sheridon's work at Parc right up through 2007.
Of course, the article says, by 2012 E-paper will be as common as napkins... It's all just around the corner.
Today's commercial uses of E-paper seem to be primarily readers, an example of which is here. While there are other uses as well, there seems to be well recognized limitations with color and speed.
But back to Mark's point: "Soon we could see the commercialization of full color and motion passive devices. That may be what is needed for epublishing to usurp traditional print publishing."
I think this is a good question.
The answer, though, is somewhat complicated and inter-twined with history and happenstance.
In terms of the history of computing and digital printing Parc is the rosetta stone, as I commented above. But only a very small number of technologies created there ever survived to be commercialized. The reason for that, among other things, is that the people that invented the technology were inventors, not business men.
Metcalfe, I think, was the chief exception, founding 3-com to commercialize Ethernet. And, even as late as the early 90's this was no sure bet. Prior to that Novell, Microsoft and others (IBM and token ring) offered networking solutions that eclipsed Ethernet. It wasn't until the Internet as we know it today took off did Ethernet's place in the world get solidified.
The metaphor at Parc was the replacement of paper for doing your work - which is not the same as replacing paper: the Alto had email, drawing programs, and so forth. But, at the end of the day, you still needed a Dover to print out the results. E-paper does not fit into this model - it was far ahead of its time in that regard. It wouldn't find a real application until maybe 2005 and later.
The early Acrobat ads focused on the same thing: Acrobat was designed to replace the need for paper on a computer. I remember watching an ad Adobe created: there was an office with copiers and typewriters. People were trying to work on documents by physically cutting and pasting and copying. People were attaching notes, marking on the paper, re-typing, re-printing and so on. Acrobat was presented as the "holy grail" that allowed allow this to mostly be done on the computer.
You also have to me what looks like simple ignorance and arrogance: Apple's iPad no Match for E-Paper.
This headline is probably true, but not in the way the authors intended it. Reading this and other E-paper ads its clear that E-paper bigots (I apologize if this offends anyone) can only imagine their product in a world where people do what they think it should be used for. In this case, behave like a book reader.
However, I think this is remarkably short sighted on their part. I have enough digital devices already - a laptop, a phone, an iPod. I don't need another one. The seem to miss the fact that the trajectory of digital products is to integrate these functions into fewer and fewer devices - not more and more specialized devices.
Do I believe that no one will see value in a Kindle? Of course not. But its a very specialized market I think - someone literally replacing the physical book with a device designed to behave as a book. But that limits E-paper to that metaphor. An iPad or laptop not only replaces the book but also does much, much more.
And finally, as I commented yesterday, the juggernaut of LCD manufacturing is just to large to be stopped or steered away. Literally there is already overcapacity in the marketplace. I talked about the cost progress of LCD's here. Its declining so rapidly that if the Virgin Space program were to be as popular the $125,000 USD cost of a flight to space will be a mere $7,500 USD in 5 years or so.
I think that E-paper will find its place in specialized applications that are well suited. My guess is that in the long run these will be manufactured or targeted applications outside the mainstream of laptop and iPad-like systems.
Print is being affect by this, not from direct replacement so much as by abandonment. Many things are still printed but the previous user base is abandoning them. Its not that an iPad user would ignore a magazine in a doctors office as short term entertainment, it just that at home that same user simply won't bother with the printed version. After doing my research I feel that E-paper is already being or is about to be abandon as well for its initial "holy grail" applications - which will leave it relegated to manufacturing and other specialized applications.
The Palo Alto Research Center (Parc), a research center for Xerox, was well known for inventing the Alto, a precursor of most modern computer systems. The Alto, as the computer was called, consisted of an networking card (Ethernet - invented by Robert Metcalfe), a bitmap display, a mouse, a custom processor, and a hard drive (2.5 Mb).
The "Dover Laser Printer" was also invented around this time (I wrote about it here). It was, as far as I know, the first networked printing device (as well as the first laser printer).
The concept of E-paper was invented by Nicholas K. Sheridon at Xerox Parc in the mid 1970's.
Some interesting material is presented here including a lengthy interview with Sheridon.
Sheridon makes an interesting point: "Much has been written about the incredible myopia of Xerox executives of the time, so I won't go into that except to say that there were numerous other opportunities to enormously expand Xerox's business that were similarly fumbled. Xerox had enough money to create an incredible research lab with top-notch people, but Xerox management could not shake off the copier mentality."
Xerox had literally invented the future of computing at Parc by 1980 or so. Everything you use and take for granted in a computing sense was created there. Supposedly Steve Jobs "stole" the idea to create the Lisa - the precursor of the Macintosh - from Parc after a visit.
But Xerox management could not understand what their research team had invented: they only understood copiers. The proof, of course, is that only the Dover was commercialized by Xerox (initially as the Xerox 9700). They eventually tried to commercialize the Alto as the Xerox Star Office - but it was a dismal failure. Even with the commercialization of the Dover as the X9700, however, the networking and so forth was discarded in favor of a mainframe channel adapter (for communicating with main frames) and an 9-track, reel-to-reel tape drive.
The link covers Gyricon and E-Ink in some detail which were various spin-offs from Sheridon's work at Parc right up through 2007.
Of course, the article says, by 2012 E-paper will be as common as napkins... It's all just around the corner.
Today's commercial uses of E-paper seem to be primarily readers, an example of which is here. While there are other uses as well, there seems to be well recognized limitations with color and speed.
But back to Mark's point: "Soon we could see the commercialization of full color and motion passive devices. That may be what is needed for epublishing to usurp traditional print publishing."
I think this is a good question.
The answer, though, is somewhat complicated and inter-twined with history and happenstance.
In terms of the history of computing and digital printing Parc is the rosetta stone, as I commented above. But only a very small number of technologies created there ever survived to be commercialized. The reason for that, among other things, is that the people that invented the technology were inventors, not business men.
Metcalfe, I think, was the chief exception, founding 3-com to commercialize Ethernet. And, even as late as the early 90's this was no sure bet. Prior to that Novell, Microsoft and others (IBM and token ring) offered networking solutions that eclipsed Ethernet. It wasn't until the Internet as we know it today took off did Ethernet's place in the world get solidified.
The metaphor at Parc was the replacement of paper for doing your work - which is not the same as replacing paper: the Alto had email, drawing programs, and so forth. But, at the end of the day, you still needed a Dover to print out the results. E-paper does not fit into this model - it was far ahead of its time in that regard. It wouldn't find a real application until maybe 2005 and later.
The early Acrobat ads focused on the same thing: Acrobat was designed to replace the need for paper on a computer. I remember watching an ad Adobe created: there was an office with copiers and typewriters. People were trying to work on documents by physically cutting and pasting and copying. People were attaching notes, marking on the paper, re-typing, re-printing and so on. Acrobat was presented as the "holy grail" that allowed allow this to mostly be done on the computer.
You also have to me what looks like simple ignorance and arrogance: Apple's iPad no Match for E-Paper.
This headline is probably true, but not in the way the authors intended it. Reading this and other E-paper ads its clear that E-paper bigots (I apologize if this offends anyone) can only imagine their product in a world where people do what they think it should be used for. In this case, behave like a book reader.
However, I think this is remarkably short sighted on their part. I have enough digital devices already - a laptop, a phone, an iPod. I don't need another one. The seem to miss the fact that the trajectory of digital products is to integrate these functions into fewer and fewer devices - not more and more specialized devices.
Do I believe that no one will see value in a Kindle? Of course not. But its a very specialized market I think - someone literally replacing the physical book with a device designed to behave as a book. But that limits E-paper to that metaphor. An iPad or laptop not only replaces the book but also does much, much more.
And finally, as I commented yesterday, the juggernaut of LCD manufacturing is just to large to be stopped or steered away. Literally there is already overcapacity in the marketplace. I talked about the cost progress of LCD's here. Its declining so rapidly that if the Virgin Space program were to be as popular the $125,000 USD cost of a flight to space will be a mere $7,500 USD in 5 years or so.
I think that E-paper will find its place in specialized applications that are well suited. My guess is that in the long run these will be manufactured or targeted applications outside the mainstream of laptop and iPad-like systems.
Print is being affect by this, not from direct replacement so much as by abandonment. Many things are still printed but the previous user base is abandoning them. Its not that an iPad user would ignore a magazine in a doctors office as short term entertainment, it just that at home that same user simply won't bother with the printed version. After doing my research I feel that E-paper is already being or is about to be abandon as well for its initial "holy grail" applications - which will leave it relegated to manufacturing and other specialized applications.
Monday, November 15, 2010
The New Publishing Systems...
I have been interested in doing away with more paper in my house. The only real paper left that comes in on a regular basis is the Wall Street Journal, a Sound-on-Sound magazine, and junk mail. At the same time I am always interested in platforms for printing (not just magazines, but books, various personal-type information like statements, bills, etc.).
The most rational choice for a reader for me would be an iPad. Most of what I read on a regular basis would work on there, e.g., the WSJ for iPad.
On the book front there are problems, though. Some books, like "The Drunkard's Walk", by Leonard Mlodinow, is available for an iPad via the free Kindle group of apps. Many are available via other forms of eReader. However, many are still not available electronically.
Now this was supposed to be the realm of Acrobat. But Acrobat peaked too early and its model is too closed (more on this below).
Recently I came across a website called www.readoz.com. This is some sort of startup publishing site that offers you the ability to publish your book, magazine or other printed work electronically. You "drop off PDF files" and they do the rest. What's interesting about this site is that they have a lot of technology to integrate your publishing with social media and they mimic the printed world very closely.
On the social media side (from this) "ReadOz digital editions also feature full search engine optimization, bookmark and share technology with 35 social networking platforms, audio/video capabilities, as well as iPod, iPad and Android applications." Their free reader applications address most of how this works.
They accept source material in PDF form (presumably plate ready but it doesn't say) organized for how the paper version of the product would be printed: You can provided belly band content, gate-fold content, and the rest. There is support of audio annotation, targeted ads, surveys, dynamic content, full ad tracking, engagement tools for email opt in, etc., and so forth (you can see the full list here).
The business model appears to be tied to the front-end - somehow the cost of using this to reduce your print-run length pays for the service. I couldn't find any details on the site regarding this however.
In terms of competitors I found www.zmags.com. This platform appears to be somewhat similar to readoz though it would appear to have been around longer and have some real customers. Their web site offers some better clues about the revenue model: "Achieve rapid ROI with digital publications. By converting just 5 percent of print subscribers to Zmags, you will save enough on print costs in one month to pay for your Zmags license for the entire year."
Zmags also offers support for marketing materials like dynamic catalogs and educational recruiting.
Educational recruiting? I have not heard of this as an industry before - and apparently its full of regulations. From the site: "A growing number of colleges and universities are now distributing digital editions of recruiting materials and media guides to prospective students. These interactive digital books offer a more dynamic way to educate prospective students and student-athletes, while complying with recruiting regulations. Zmags, the industry leader in interactive digital publishing software, enables universities to provide prospective students with a high-quality interactive reading experience that includes digital pictures, video clips, live links, news feeds, and more, which fully immerses prospective students into their institutions."
Now the Sound-on-Sound magazine that I subscribe to uses some kind of technology like this - but I cannot tell if its home grown or from a service like one of these. So far I have not been pleased with the electronic version of this - its kind of clunky to have to zoom and pan around to see the entire page at any given point. (I like to view the entire page - lots of articles in technical magazines have things on the spreads that cross-reference each other, e.g., a box with Pros/Cons and pricing on the opposite page from the article intro.
I don't see the current iPad models making this any better (I have 17" Mac laptops which are also clunky for this) and this is the primary reason I would not switch.
Also, I have to wonder about sharing a lot of content via Facebook and so forth with these platforms. While I can see how this might be handy on occasion in general I don't think people on your Facebook will want to dive into too much detail. I think that the length of your posts on Facebook correspond to your age - the younger you are the shorter the posts. Linking to long-winded articles probably won't do much for the younger set.
Overall, though, I cannot see how publishing would not move in this direction - whether with these particular tools I described or with others - in any case the die has been cast. I think that one thing that will be needed to succeed will be a 17" display - like the one on the MacBook - but without a keyboard and turned 90-degrees. Ideally I'd like to see double that - almost like two 17" displays. (Perhaps a version with which you could fold the keyboard all the way back around the display would do it.) The iPad is the first step - but I think its not quite enough (sorry Steve).
Part of this too is that I am a geezer and like larger type. My kids and grandkids have no problem with tiny text on tiny displays - me, I don't care for it. Though I can see that texting on a 17" iPad might be distracting while driving.
Acrobat wants to do all of this but it can't. I think the reason is that its too closely tied to the publishing end (in terms of creation) and not properly tied to reading end. It supports everything all the rest do - interactivity and so on - but its not quite the same. I think its also a bit too "technical" for a lot of things - particularly basic reading. There are a lot more computer users today with a lot less knowledge of publishing and associated issues.
I think Flash is also part of the problem - its not really integrated with Acrobat and yet most of the effects like page turning and such that you see with eReaders and such are Flash-like. Adobe, I think, will remain a player as long as print is involved. But once the scale tips away from print, i.e., 5% read the magazine in print instead of 5% electronically, the publishing baggage (CMYK, plate-layouts, etc.) will be cast off and replaced with newer, sleeker tools.
The most rational choice for a reader for me would be an iPad. Most of what I read on a regular basis would work on there, e.g., the WSJ for iPad.
On the book front there are problems, though. Some books, like "The Drunkard's Walk", by Leonard Mlodinow, is available for an iPad via the free Kindle group of apps. Many are available via other forms of eReader. However, many are still not available electronically.
Now this was supposed to be the realm of Acrobat. But Acrobat peaked too early and its model is too closed (more on this below).
Recently I came across a website called www.readoz.com. This is some sort of startup publishing site that offers you the ability to publish your book, magazine or other printed work electronically. You "drop off PDF files" and they do the rest. What's interesting about this site is that they have a lot of technology to integrate your publishing with social media and they mimic the printed world very closely.
On the social media side (from this) "ReadOz digital editions also feature full search engine optimization, bookmark and share technology with 35 social networking platforms, audio/video capabilities, as well as iPod, iPad and Android applications." Their free reader applications address most of how this works.
They accept source material in PDF form (presumably plate ready but it doesn't say) organized for how the paper version of the product would be printed: You can provided belly band content, gate-fold content, and the rest. There is support of audio annotation, targeted ads, surveys, dynamic content, full ad tracking, engagement tools for email opt in, etc., and so forth (you can see the full list here).
The business model appears to be tied to the front-end - somehow the cost of using this to reduce your print-run length pays for the service. I couldn't find any details on the site regarding this however.
In terms of competitors I found www.zmags.com. This platform appears to be somewhat similar to readoz though it would appear to have been around longer and have some real customers. Their web site offers some better clues about the revenue model: "Achieve rapid ROI with digital publications. By converting just 5 percent of print subscribers to Zmags, you will save enough on print costs in one month to pay for your Zmags license for the entire year."
Zmags also offers support for marketing materials like dynamic catalogs and educational recruiting.
Educational recruiting? I have not heard of this as an industry before - and apparently its full of regulations. From the site: "A growing number of colleges and universities are now distributing digital editions of recruiting materials and media guides to prospective students. These interactive digital books offer a more dynamic way to educate prospective students and student-athletes, while complying with recruiting regulations. Zmags, the industry leader in interactive digital publishing software, enables universities to provide prospective students with a high-quality interactive reading experience that includes digital pictures, video clips, live links, news feeds, and more, which fully immerses prospective students into their institutions."
Now the Sound-on-Sound magazine that I subscribe to uses some kind of technology like this - but I cannot tell if its home grown or from a service like one of these. So far I have not been pleased with the electronic version of this - its kind of clunky to have to zoom and pan around to see the entire page at any given point. (I like to view the entire page - lots of articles in technical magazines have things on the spreads that cross-reference each other, e.g., a box with Pros/Cons and pricing on the opposite page from the article intro.
I don't see the current iPad models making this any better (I have 17" Mac laptops which are also clunky for this) and this is the primary reason I would not switch.
Also, I have to wonder about sharing a lot of content via Facebook and so forth with these platforms. While I can see how this might be handy on occasion in general I don't think people on your Facebook will want to dive into too much detail. I think that the length of your posts on Facebook correspond to your age - the younger you are the shorter the posts. Linking to long-winded articles probably won't do much for the younger set.
Overall, though, I cannot see how publishing would not move in this direction - whether with these particular tools I described or with others - in any case the die has been cast. I think that one thing that will be needed to succeed will be a 17" display - like the one on the MacBook - but without a keyboard and turned 90-degrees. Ideally I'd like to see double that - almost like two 17" displays. (Perhaps a version with which you could fold the keyboard all the way back around the display would do it.) The iPad is the first step - but I think its not quite enough (sorry Steve).
Part of this too is that I am a geezer and like larger type. My kids and grandkids have no problem with tiny text on tiny displays - me, I don't care for it. Though I can see that texting on a 17" iPad might be distracting while driving.
Acrobat wants to do all of this but it can't. I think the reason is that its too closely tied to the publishing end (in terms of creation) and not properly tied to reading end. It supports everything all the rest do - interactivity and so on - but its not quite the same. I think its also a bit too "technical" for a lot of things - particularly basic reading. There are a lot more computer users today with a lot less knowledge of publishing and associated issues.
I think Flash is also part of the problem - its not really integrated with Acrobat and yet most of the effects like page turning and such that you see with eReaders and such are Flash-like. Adobe, I think, will remain a player as long as print is involved. But once the scale tips away from print, i.e., 5% read the magazine in print instead of 5% electronically, the publishing baggage (CMYK, plate-layouts, etc.) will be cast off and replaced with newer, sleeker tools.
Friday, November 12, 2010
Google: What Would Jobs Have Done?
I was reading this article at CIO regarding how Google might look a little different had Steve Jobs taken the helm in place of Eric Schmidt. I bring this up because a while back I wrote "Whistling Past the Graveyard" which talked about why companies that get defocused from their primary expertise tend to get lost. Google is another case in point.
Google defined the concept of a "search company". There were lots of search technologies before Google, and many that have come after, like Bing. But Google remains the master at search.
Unfortunately, this is no longer the chief focus as far as I can see.
So what about the idea of Steve Jobs in charge at Google? I think that would have made Google far more focused on its core search technology and much less focused on satellite things like blogspot, gmail, chrome, android, cell phones, wave, the supposed Google Me Facebook replacement and all the rest.
Not many of these things are really winners - particularly from the "big picture" perspective.
I would also say this is exactly why Steve Jobs would not have taken the head job at Google. The core search engine model for Google, while technologically exciting, is far from an ego and marketing pizazz blockbuster and it really doesn't offer much for Jobs. Can you really imagine him standing up on the stage a Moscone Center talking about some low-level Python search wizardry? (Yes, he could pull it off, but it wouldn't be the same as the Apple events he hosts.)
Jobs likes to focus on making things great. He does this by grinding away the superfluous crap that make things annoying and complicated and getting down the real nub of the issue, e.g., the iPod interface. The Google search screen is already this way - as are the returned results (at least for the most part). What would he be left to focus on?
Making clunky weird things like wave and chrome.
As wonderful as Google is today at what they do there are, however, still search problems to be solved. My concern with Google these days is that they have become so focused on the other things I mentioned that making search work right is no longer the number one priority.
Do I need really Chrome? Does my browser really need to be that much faster?
Android has its own set of issues, particularly about not being as "open" as Google claims.
I have also written here about Google's bizarre ideas about "cloud printing". HP has scooped them in this regard - just watch their TV ads.
Certainly I use blogspot at Google for this blog. It works okay though at some point I may consider moving the elsewhere to have more control over it.
Many people love gmail. Now I really have a hard time understanding why this is so. People put personal things in email, business things, things they say that are private. Why would you trust Google with this information? Particularly as the company grows and grows beyond all real means to control it.
I fear the insidious nature of Google's approach to content. Particularly my content (not me personally necessarily, but the whole notion of people putting their private information into Google). Now this blog is a public blog - and "behind the scenes" there isn't really anything - I don't keep any email on Google, no secret links or stores of data, and I do most of my research outside of blogspot. So while Google hosts my blog it doesn't have any personal information about me or my blog or, for that matter, my company.
And Google tends to create weird standards for things, like mail attachments. We have a customer that loved to send gmail's with these bizarre Google attachments that would not work on a Mac.
How nice. I guess no on there uses a Mac.
Google now finds itself in the position where its no longer the "technological hot ticket". Eric Schmidt recently decided to give all the employees a 10% raise. I doubt very much this was out of the goodness of his cold black heart. He's doing this to keep people. Keep people for escaping to Facebook and other startups.
This is going to make these defocused activities harder to support because most of them are pet projects cooked up by employees in the first place.
The problem with companies and privacy as they get larger and larger is that not everyone follows the privacy code the founders established. Eventually there will be unhappy employees who leave - taking secure information with them. And when they go who's gmail passwords will go with them?
"Don't be evil" only gets you so far - there's a lot of latitude around "evil".
Things like Facebook are now starting to breath down Google's neck. Facebook is more exciting and starting to draw employees away from Google. This will make Google's life harder and put more stress on the whole mess.
So, for me, what does all this mean?
First off, I think Google needs to stick to search and fix what's already broken with search.
My biggest beef has always been, being in the PDF business, that you cannot effectively search for things about PDF without getting a list of things that are PDF, i.e., PDF files about the things you are interested in.
For example, suppose I am interested in whether there's a PDF viewer on a specific kind of microprocessor, e.g., the ARM LPC1343. You would think that Googling "LPC1343 PDF" might work - but it doesn't. Instead I get a bunch of PDF files about the LPC1343.
This is just an example, but I find this quite annoying and not helpful at all.
So Google, if your listening, get back on track at what you do best - before you too are "Whistling Past the Graveyard".
Google defined the concept of a "search company". There were lots of search technologies before Google, and many that have come after, like Bing. But Google remains the master at search.
Unfortunately, this is no longer the chief focus as far as I can see.
So what about the idea of Steve Jobs in charge at Google? I think that would have made Google far more focused on its core search technology and much less focused on satellite things like blogspot, gmail, chrome, android, cell phones, wave, the supposed Google Me Facebook replacement and all the rest.
Not many of these things are really winners - particularly from the "big picture" perspective.
I would also say this is exactly why Steve Jobs would not have taken the head job at Google. The core search engine model for Google, while technologically exciting, is far from an ego and marketing pizazz blockbuster and it really doesn't offer much for Jobs. Can you really imagine him standing up on the stage a Moscone Center talking about some low-level Python search wizardry? (Yes, he could pull it off, but it wouldn't be the same as the Apple events he hosts.)
Jobs likes to focus on making things great. He does this by grinding away the superfluous crap that make things annoying and complicated and getting down the real nub of the issue, e.g., the iPod interface. The Google search screen is already this way - as are the returned results (at least for the most part). What would he be left to focus on?
Making clunky weird things like wave and chrome.
As wonderful as Google is today at what they do there are, however, still search problems to be solved. My concern with Google these days is that they have become so focused on the other things I mentioned that making search work right is no longer the number one priority.
Do I need really Chrome? Does my browser really need to be that much faster?
Android has its own set of issues, particularly about not being as "open" as Google claims.
I have also written here about Google's bizarre ideas about "cloud printing". HP has scooped them in this regard - just watch their TV ads.
Certainly I use blogspot at Google for this blog. It works okay though at some point I may consider moving the elsewhere to have more control over it.
Many people love gmail. Now I really have a hard time understanding why this is so. People put personal things in email, business things, things they say that are private. Why would you trust Google with this information? Particularly as the company grows and grows beyond all real means to control it.
I fear the insidious nature of Google's approach to content. Particularly my content (not me personally necessarily, but the whole notion of people putting their private information into Google). Now this blog is a public blog - and "behind the scenes" there isn't really anything - I don't keep any email on Google, no secret links or stores of data, and I do most of my research outside of blogspot. So while Google hosts my blog it doesn't have any personal information about me or my blog or, for that matter, my company.
And Google tends to create weird standards for things, like mail attachments. We have a customer that loved to send gmail's with these bizarre Google attachments that would not work on a Mac.
How nice. I guess no on there uses a Mac.
Google now finds itself in the position where its no longer the "technological hot ticket". Eric Schmidt recently decided to give all the employees a 10% raise. I doubt very much this was out of the goodness of his cold black heart. He's doing this to keep people. Keep people for escaping to Facebook and other startups.
This is going to make these defocused activities harder to support because most of them are pet projects cooked up by employees in the first place.
The problem with companies and privacy as they get larger and larger is that not everyone follows the privacy code the founders established. Eventually there will be unhappy employees who leave - taking secure information with them. And when they go who's gmail passwords will go with them?
"Don't be evil" only gets you so far - there's a lot of latitude around "evil".
Things like Facebook are now starting to breath down Google's neck. Facebook is more exciting and starting to draw employees away from Google. This will make Google's life harder and put more stress on the whole mess.
So, for me, what does all this mean?
First off, I think Google needs to stick to search and fix what's already broken with search.
My biggest beef has always been, being in the PDF business, that you cannot effectively search for things about PDF without getting a list of things that are PDF, i.e., PDF files about the things you are interested in.
For example, suppose I am interested in whether there's a PDF viewer on a specific kind of microprocessor, e.g., the ARM LPC1343. You would think that Googling "LPC1343 PDF" might work - but it doesn't. Instead I get a bunch of PDF files about the LPC1343.
This is just an example, but I find this quite annoying and not helpful at all.
So Google, if your listening, get back on track at what you do best - before you too are "Whistling Past the Graveyard".
Subscribe to:
Posts (Atom)

