Big Data Basics - Part 1 - Introduction to Big Data

By:   |   Comments (16)   |   Related: More > Big Data


Problem

I have been hearing the term Big Data for a while now and would like to know more about it. Can you explain what this term means, how it evolved, and how we identify Big Data and any other relevant details?

Solution

Big Data has been a buzz word for quite some time now and it is catching popularity faster than pretty much anything else in the technology world. In this tip, let us understand what this buzz word is all about, what is its significance, why you should care about it, and more.

What is Big Data?

Wikipedia defines "Big Data" as a collection of data sets so large and complex that it becomes difficult to process using on-hand database management tools or traditional data processing applications.

In simple terms, "Big Data" consists of very large volumes of heterogeneous data that is being generated, often, at high speeds.  These data sets cannot be managed and processed using traditional data management tools and applications at hand.  Big Data requires the use of a new set of tools, applications and frameworks to process and manage the data.

Evolution of Data / Big Data

Data has always been around and there has always been a need for storage, processing, and management of data, since the beginning of human civilization and human societies. However, the amount and type of data captured, stored, processed, and managed depended then and even now on various factors including the necessity felt by humans, available tools/technologies for storage, processing, management, effort/cost, ability to gain insights into the data, make decisions, and so on.

Going back a few centuries, in the ancient days, humans used very primitive ways of capturing/storing data like carving on stones, metal sheets, wood, etc. Then with new inventions and advancements a few centuries in time, humans started capturing the data on paper, cloth, etc. As time progressed, the medium of capturing/storage/management became punching cards followed by magnetic drums, laser disks, floppy disks, magnetic tapes, and finally today we are storing data on various devices like USB Drives, Compact Discs, Hard Drives, etc.

In fact the curiosity to capture, store, and process the data has enabled human beings to pass on knowledge and research from one generation to the next, so that the next generation does not have to re-invent the wheel.

As we can clearly see from this trend, the capacity of data storage has been increasing exponentially, and today with the availability of the cloud infrastructure, potentially one can store unlimited amounts of data. Today Terabytes and Petabytes of data is being generated, captured, processed, stored, and managed.

Characteristics of Big Data - The Three V's of Big Data

When do we say we are dealing with Big Data? For some people 1TB might seem big, for others 10TB might be big, for others 100GB might be big, and something else for others. This term is qualitative and it cannot really be quantified. Hence we identify Big Data by a few characteristics which are specific to Big Data. These characteristics of Big Data are popularly known as Three V's of Big Data.

The three v's of Big Data are Volume, Velocity, and Variety as shown below.

Characteristics of Big Data - The Three V's of Big Data

Volume

Volume refers to the size of data that we are working with. With the advancement of technology and with the invention of social media, the amount of data is growing very rapidly.  This data is spread across different places, in different formats, in large volumes ranging from Gigabytes to Terabytes, Petabytes, and even more. Today, the data is not only generated by humans, but large amounts of data is being generated by machines and it surpasses human generated data. This size aspect of data is referred to as Volume in the Big Data world.

Velocity

Velocity refers to the speed at which the data is being generated. Different applications have different latency requirements and in today's competitive world, decision makers want the necessary data/information in the least amount of time as possible.  Generally, in near real time or real time in certain scenarios. In different fields and different areas of technology, we see data getting generated at different speeds. A few examples include trading/stock exchange data, tweets on Twitter, status updates/likes/shares on Facebook, and many others. This speed aspect of data generation is referred to as Velocity in the Big Data world.

Variety

Variety refers to the different formats in which the data is being generated/stored. Different applications generate/store the data in different formats. In today's world, there are large volumes of unstructured data being generated apart from the structured data getting generated in enterprises. Until the advancements in Big Data technologies, the industry didn't have any powerful and reliable tools/technologies which can work with such voluminous unstructured data that we see today. In today's world, organizations not only need to rely on the structured data from enterprise databases/warehouses, they are also forced to consume lots of data that is being generated both inside and outside of the enterprise like clickstream data, social media, etc. to stay competitive. Apart from the traditional flat files, spreadsheets, relational databases etc., we have a lot of unstructured data stored in the form of images, audio files, video files, web logs, sensor data, and many others. This aspect of varied data formats is referred to as Variety in the Big Data world.

Sources of Big Data

Just like the data storage formats have evolved, the sources of data have also evolved and are ever expanding.  There is a need for storing the data into a wide variety of formats. With the evolution and advancement of technology, the amount of data that is being generated is ever increasing. Sources of Big Data can be broadly classified into six different categories as shown below.

Sources of Big Data

Enterprise Data

There are large volumes of data in enterprises in different formats. Common formats include flat files, emails, Word documents, spreadsheets, presentations, HTML pages/documents, pdf documents, XMLs, legacy formats, etc. This data that is spread across the organization in different formats is referred to as Enterprise Data.

Transactional Data

Every enterprise has some kind of applications which involve performing different kinds of transactions like Web Applications, Mobile Applications, CRM Systems, and many more. To support the transactions in these applications, there are usually one or more relational databases as a backend infrastructure. This is mostly structured data and is referred to as Transactional Data.

Social Media

This is self-explanatory. There is a large amount of data getting generated on social networks like Twitter, Facebook, etc. The social networks usually involve mostly unstructured data formats which includes text, images, audio, videos, etc. This category of data source is referred to as Social Media.

Activity Generated

There is a large amount of data being generated by machines which surpasses the data volume generated by humans. These include data from medical devices, censor data, surveillance videos, satellites, cell phone towers, industrial machinery, and other data generated mostly by machines. These types of data are referred to as Activity Generated data.

Public Data

This data includes data that is publicly available like data published by governments, research data published by research institutes, data from weather and meteorological departments, census data, Wikipedia, sample open source data feeds, and other data which is freely available to the public. This type of publicly accessible data is referred to as Public Data.

Archives

Organizations archive a lot of data which is either not required anymore or is very rarely required. In today's world, with hardware getting cheaper, no organization wants to discard any data, they want to capture and store as much data as possible. Other data that is archived includes scanned documents, scanned copies of agreements, records of ex-employees/completed projects, banking transactions older than the compliance regulations.  This type of data, which is less frequently accessed, is referred to as Archive Data.

Formats of Data

Data exists in multiple different formats and the data formats can be broadly classified into two categories - Structured Data and Unstructured Data.

Structured data refers to the data which has a pre-defined data model/schema/structure and is often either relational in nature or is closely resembling a relational model. Structured data can be easily managed and consumed using the traditional tools/techniques. Unstructured data on the other hand is the data which does not have a well-defined data model or does not fit well into the relational world.

Structured data includes data in the relational databases, data from CRM systems, XML files etc. Unstructured data includes flat files, spreadsheets, Word documents, emails, images, audio files, video files, feeds, PDF files, scanned documents, etc.

Big Data Statistics

  • 100 Terabytes of data is uploaded to Facebook every day
  • Facebook Stores, Processes, and Analyzes more than 30 Petabytes of user generated data
  • Twitter generates 12 Terabytes of data every day
  • LinkedIn processes and mines Petabytes of user data to power the "People You May Know" feature
  • YouTube users upload 48 hours of new video content every minute of the day
  • Decoding of the human genome used to take 10 years. Now it can be done in 7 days
  • 500+ new websites are created every minute of the day

Source: Wikibon - A Comprehensive List of Big Data Statistics

In this tip we were introduced to Big Data, how it evolved, what are its primary characteristics, what are the sources of data, and a few statistics showing how large volumes of heterogeneous data is being generated at different speeds.

References
Next Steps
  • Explore more about Big Data.  Do some of your own searches to see what you can find.
  • Stay tuned for future tips in this series to learn more about the Big Data ecosystem.


sql server categories

sql server webinars

subscribe to mssqltips

sql server tutorials

sql server white papers

next tip



About the author
MSSQLTips author Dattatrey Sindol Dattatrey Sindol has 8+ years of experience working with SQL Server BI, Power BI, Microsoft Azure, Azure HDInsight and more.

This author pledges the content of this article is based on professional experience and not AI generated.

View all my tips



Comments For This Article




Thursday, April 11, 2019 - 6:31:55 AM - Prakesh Back To Top (79528)

Really nice article. I would like to know, as the series of article related to big data  were written back in 2013, are there any changes since past 6 years.


Tuesday, July 11, 2017 - 10:50:06 PM - Ming Back To Top (59268)

 Hi Datta, 

Thanks, Great education on the Big Data nd the basic architecture. Since the article 2014 to now 2017,

1. have there been a lot of evolution or changes in the architecture ( semantic layer, atomic layer ) as i see that with the huge data consumtion and age of IOT. 

2. I also see there is folks that like Hadoop ( ie. Hive) vs mpp ( ie Redshift etc), can you share your thoughts on the pro and cons. 

 Thanks for the education and hope to learn more. 

Ming


Wednesday, May 10, 2017 - 3:21:39 AM - pvrao Back To Top (55643)

 Very informative as I'm looking to get into this for futher steps, Thanks for sharing this in such simple terms. 

 


Wednesday, May 25, 2016 - 8:24:18 AM - NitinPune Back To Top (41556)

 This page, i believe, is the kick of point for knowing about Big Data. Thanks for sharing this in such simple terms. 

 


Monday, October 19, 2015 - 1:55:21 PM - ed w Back To Top (38934)

Very informative as I'm looking to get into this and out of construction.....thanks for posting


Wednesday, September 23, 2015 - 9:10:35 AM - Gopinath Back To Top (38733)

Hi Datta,

 This kind of information is what i was looking for .

Very informative and easy to understand.Thanks a lot !!!

 

 


Wednesday, December 24, 2014 - 10:22:20 AM - Jay Back To Top (35755)

Hey Datta

 

Thanks a lot!! 

Useful information to understand the basics of "Big-Data".


Monday, September 22, 2014 - 2:01:42 PM - Srinivas Back To Top (34659)

 

Useful article to start learnig Big data. Thanks

 

Srini

 


Tuesday, September 9, 2014 - 12:33:58 PM - vivek sharma Back To Top (34455)

Good introduction of Big data.


Friday, June 20, 2014 - 8:58:39 AM - Sankar Reddy Back To Top (32324)

Good overview about bigdata..


Tuesday, March 4, 2014 - 2:56:31 PM - Jeremy Kadlec Back To Top (29639)

Daniel and Datta,

This has been updated.

Thank you,
Jeremy Kadlec
MSSQLTips.com Community Co-Leader


Tuesday, March 4, 2014 - 10:15:13 AM - Dattatrey Sindol (Datta) Back To Top (29635)

Hi Daniel,

Yes, it is a typo. Thank you for pointing it out.

We will get it corrected.

Best Regards,

Dattatrey Sindol (Datta)


Tuesday, March 4, 2014 - 9:28:03 AM - Daniel DeLuna Back To Top (29634)

Thanks for the great article.  I don't understand the use of the term "censor" data.  Did you mean "sensor" data.  This is a minor point and I'm just looking for clarification.


Monday, February 3, 2014 - 8:49:59 AM - KRISHNA KANT Back To Top (29316)

It is good conceptual overview about the BIG DATA....THANKS.


Thursday, January 30, 2014 - 12:17:47 AM - Yogesh Back To Top (29276)

Awesome ! I just want to start in this technology......

Can you please give idea how long it take to learn Big Data and cloud.


Sunday, December 29, 2013 - 8:22:37 AM - Prasad Back To Top (27905)

Good start!! 















get free sql tips
agree to terms