Three cheers for Big Data on Cloud
This week MapR announced that it broke the MinuteSort record, where it sorted 15 billion 100 byte records in less than a minute. This calls for a celebration. We are generating data at an enormous pace. This will continue to grow exponentially, as more and more devices come online and get connected through the Internet of Things. The real problem will be making sense of the data at a reasonable entry and ongoing price.
The previous record was set by Microsoft’s specialized software using private machine clusters that needed 27,000 cores. This record was set up a cluster of 4206 cores, built from Google’s Compute Cluster that will soon be commercially available, and minor tweaking of MapR’s distribution of Open Source Hadoop distribution.
Big Data processing on cloud is improving fast. Last year, two friends, after having built and run Facebook’s data service, have set up Qubole, a cloud data platform. All this investment and startup activity will hopefully allow us to unlock large number of practical use cases from the big data, and foster further innovation













