MIT Rethinks Big Data Processing

Research by a small team at the Massachusetts Institute of Technology may turn out to help streamline the processing of big data--those terabytes of streaming data that are generated from GPSs in smartphones and a multitude of other sensors. The basic idea is to create "succinct representations" of huge data sets so that existing algorithms can handle them more efficiently.

As described in "The Single Pixel GPS: Learning Big Data Signals from Tiny Coresets," a paper presented at the Association for Computing Machinery's International Conference on Advances in Geographic Information Systems, three MIT researchers have figured out how to represent data so that it takes up less space in memory while still being processed in conventional ways. That's useful because it means the technique can be used with existing algorithms rather than having to replace them with new ones.

The researchers applied the technique to the processing of two-dimensional location data generated by GPS receivers. According to Daniela Rus, a professor of computer science and engineering and director of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL), these receivers take position readings every 10 seconds. That adds up to about a gigabyte of data each day. Systems that attempt to analyze traffic patterns from readings sent by a massive number of cars can easily be bogged down by the volume of data generated.

What the scientists have figured out is that the analysis doesn't need to encompass each point of data generated by a given car--only some of it, such as when the car is turning. The path between that point and the next turn could be approximated by a straight line. The collection of those sets of data form a new "coreset" that can be compressed on the run, as it were.

The researchers' algorithm has to find a series of line segments that most accurately defines the data points. The algorithm also stores the exact coordinates of a random sampling of the points, which stand in for the potential randomness of the unsampled points in the calculations.

The technique, which encompasses a great deal of mathematics, is a tradeoff between "accuracy and complexity," said Dan Feldman, a post-doctoral student in Rus' group and lead author on the new paper. It's the combination of linear estimates and random sampling that allows the algorithm to compress data in chunks; as new data arrives, the algorithm does recalculations.

What's the point? For all practical purposes, many potential uses for big data don't stand up to the processing they would require. The MIT team's approach suggests that a slightly erroneous approximation is better than a calculation that doesn't get performed at all. Now the scientists must consider uses for the technique that have similar characteristics to the use of GPS receiver data.

One application under consideration by Feldman is the analysis of video data. Each scene might be considered comparable to a line segment; the shift from one scene to another is like the car turning. And sample frames from a scene could provide that random sampling.

This isn't the only research being done on campus in the area of big data. In May 2012 MIT was selected to host "bigdata@CSAIL," a new Intel-sponsored research center focused on developing techniques for working with big data.

About the Author

Dian Schaffhauser is a former senior contributing editor for 1105 Media's education publications THE Journal, Campus Technology and Spaces4Learning.

Featured

  • CIS Sandbox Tutor Eli Blouin at Bentley University imagines a future technology learning partner that supports curiosity, critical thinking, feedback, and personalized learning.

    Student Voices: Technology as a Future Learning Partner

    A data analytics and marketing major at Bentley University and a distinguished lecturer at the university explore the future of "technology learning partners" — technologies that participate in the learning process by asking questions and giving feedback, not just information search results.

  • Abstract energetic glowing lines with particles

    Meta Steps Up Enterprise AI Ambitions with Muse Spark Launch

    Meta has announced the launch of Muse Spark 1.1, a multimodal reasoning model designed for agentic AI, alongside a new Meta Model API that gives developers access to the model for the first time.

  • Businessman using laptop analyzing data and growth graph chart

    AI Budgets in Education Show No Sign of Decline

    The vast majority of education organizations (98%) expect their AI infrastructure budgets to either increase or hold steady over the next year, according to a recent report from cloud storage provider Wasabi.

  • Dana Brunson facilitates a roundtable discussion with research and higher education IT leaders

    Internet2: Closing the Access Gap for Research Cyberinfrastructure

    Internet2's Research Engagement Team brings CIOs and other campus technology leadership together with research computing and data facilitators, forming a community that enables research cyberinfrastructure at institutions of all types and sizes.