Showing posts with label smart data. Show all posts
Showing posts with label smart data. Show all posts

Monday, June 09, 2014

What is Big Data and when will it be Smart Data?

Big data is cell phone users having an average of 100 interactions with their phone per day, all of which generate computerized records (100s of trillions of records). Big data is every financial market transaction, every passenger on every airplane flight, every shipped container, every transportation conveyance, every tweet, and every Internet post (all in the 100s of billions or trillions of records). Every transaction for all time.

One area of long-standing data interest is mortgage statistics since mis-estimating prepayments can cost investors billions of dollars. This raises the question of how prepayment risk is still being mis-estimated. Irrespective, mortgage data is one of the fastest growing kinds of data, both by row and column of tracked data, growing at more than 2x Moore’s law on a log chart (Moore’s law reflects the hardware on which the data is stored and manipulated (algorithms somewhat fill the gap)). This begs the question of smart data rather than big data.

There is much talk about all types of data growing (and data scientists being the biggest category of job growth), but the size of big data should surely be one of its most basic attributes. What is much more relevant is the value that big data provides through its use. For example, how has having more rows and columns in mortgage-tracking spreadsheets improved (if at all) prepayment prediction?

Like genomics, many big data problems are in the early stages of ‘the diffs,’ not knowing which part of the data is salient to keep out of the 99% that may be useless. ‘The diffs’ are the differences, the differences between a sample data set and the reference/normal data set that constitute salience and allow the rest of the data to be discarded.

Sunday, February 02, 2014

Turning Big Data into Smart Data

A key contemporary trend is big data - the creation and manipulation of large complex data sets that must be stored and managed in the cloud as they are too unwieldy for local computers. Big data creation is currently on the order of zettabytes (10007 bytes) per year, in roughly equal amounts by four segments: individuals (photos, video), companies (transaction monitoring), governments (surveillance (e.g.; the new Utah Data Center)), and scientific research (astronomical observations).

Big data fanfare abounds, we continuously hear announcements like more data was created last year than in the entire history of humanity, and that data creation is on a two year-doubling cycle. Better cheap fast storage has been the historical answer to supporting the ever-growing capacity to generate data, however this is not necessarily the best solution. Already much collected data is thrown away (e.g.; CCTV footage, real-time surgery video, and genome sequencing data) without saving anything. Much of stored data remains unused, and not cleaned up into a form that is human-usable since this is costly and challenging (de-duplication a primary example).

Turning big data into smart data means moving away from data fundamentalism, the idea that data must be collected, and that data collection in itself is an ends rather than a means. Advancement comes from smart data, not more data; being able to cleanly extract and use salient aspects of data (e.g.; the ‘diffs,’ for example identifying relevant genomic polymorphisms from the whole genome sequence), not just generate and discard or mindlessly store.