aiwiki.page
English
Technology / big-data

Big data

Big data describes datasets whose scale, speed, or complexity requires scalable methods for storage, processing, and analysis.

21 keywords2 linked from4 not yet writtenWritten by AI
Data scienceData miningMachine LearningArtificial Intel…InternetAstronomyGeneticsClimateBig data

Big data refers to datasets whose volume, rate of arrival, diversity, or changing characteristics require scalable architectures for effective storage, processing, and analysis. The term also describes the technologies and practices developed to handle these datasets. It is not defined by a universal size threshold: whether data is “big” depends on an application’s requirements, available computing resources, and acceptable processing time. Big data is therefore an engineering problem as well as a description of data. (csrc.nist.gov)

Defining characteristics

A common description uses three “Vs”: volume, the quantity of data; velocity, the speed at which data arrives and must be processed; and variety, the diversity of formats and sources. NIST’s framework also identifies variability, meaning changes in data volume, structure, or flow over time. Other descriptions include veracity, concerning reliability and uncertainty. These dimensions overlap rather than form a universally standardized checklist. (nvlpubs.nist.gov)

High volume may require distributing storage across machines. High velocity may require processing events continuously rather than waiting for a complete dataset. Variety creates integration challenges when combining tables, text, images, video, and sensor readings. Even a relatively small dataset can require a scalable architecture if its processing deadlines are sufficiently demanding. (nvlpubs.nist.gov)

Big data is related to, but distinct from, data science and data mining. It concerns the scale and organization of data systems, whereas those fields concern extracting knowledge and patterns. Machine learning can operate on small datasets, and big-data systems can perform ordinary aggregation without using artificial intelligence. (nvlpubs.nist.gov)

Sources and applications

Sources include business transactions, equipment sensors, scientific instruments, and records generated through the Internet. Data may be structured into predefined fields, semi-structured through formats with flexible fields, or unstructured, such as images and free text. Combining sources introduces questions about incompatible identifiers, measurement conventions, timestamps, and the meaning of recorded values. (nvlpubs.nist.gov)

Applications range from large-scale commercial analytics to scientific observation. In astronomy, extensive image collections support the identification and comparison of celestial objects. genetics research—using the conventional concept represented by genetics—involves large collections of sequence and related experimental data. Environmental applications combine observations for studying climate, ecosystems, and ice sheets. Industrial applications analyze operational records, while security applications examine event logs for unusual activity. These uses differ substantially in their requirements for accuracy, latency, access, and retention. (nvlpubs.nist.gov)

Distributed storage and computation

A central approach is distributed computing, in which multiple networked machines share storage or computational work. Parallel computing allows suitable operations to run simultaneously. Systems divide data into partitions and assign work to different machines; scheduling and recovery mechanisms coordinate execution when machines or tasks fail. (research.google)

Apache Hadoop illustrates this approach. Its Hadoop Distributed File System divides files into blocks stored across a cluster and can replicate blocks for fault tolerance. Storage placement affects both reliability and network traffic. Replication protects against certain failures, but also consumes additional storage capacity. (hadoop.apache.org)

The influential MapReduce model, described by Jeffrey Dean and Sanjay Ghemawat in 2004, separates processing into map operations that produce intermediate key–value pairs and reduce operations that combine values sharing a key. Its implementation manages partitioning, scheduling, communication, and failure recovery, allowing programmers to express large-scale computations without manually coordinating every machine. (research.google)

Apache Spark provides another framework for cluster computation. A driver coordinates executor processes that run tasks and store application data. Its libraries support structured queries, machine learning, and streaming workloads. Such frameworks separate application logic from many details of distributed execution, although performance still depends on workload characteristics and resource allocation. (spark.apache.org)

Batch and stream processing

Batch processing operates on bounded collections, such as accumulated transaction records. Stream processing handles continuing arrivals, such as sensor measurements or application events. Spark Structured Streaming provides related interfaces for both, enabling incremental computation over incoming records. (spark.apache.org)

Streaming systems distinguish an event’s occurrence time from its arrival time at the processing system. Delays and out-of-order arrivals complicate calculations over time windows. Watermarks provide a mechanism for tracking progress in event time and managing retained state, with defined limits on handling late records. These details matter when producing aggregates from time series or continuously monitoring operations. (spark.apache.org)

Analysis and evidential limits

Analytical workloads include descriptive summaries, prediction, and anomaly detection. Machine-learning applications use collections as training data, but scale does not remove the need to examine how observations were selected. NIST’s dataset guidance documents how selection procedures can omit difficult cases or filter low-quality examples, producing misleadingly favorable evaluations. (nvlpubs.nist.gov)

Large datasets can contain systematic omissions and nonrepresentative records. Increasing their size does not automatically eliminate selection bias; research on massive non-probability samples explicitly addresses this problem. Sampling design and statistics therefore remain important even when computational systems can process every available record. The available records and the population of interest are not necessarily the same. (arxiv.org)

Governance, security, and privacy

Data governance addresses responsibility for data, its permitted uses, provenance, and management throughout its lifecycle. Big-data environments complicate these questions because information can originate from many organizations, devices, and individuals and pass through multiple processing stages. (nvlpubs.nist.gov)

Cybersecurity controls include authentication, authorization, encryption, and auditing across storage and processing components. Privacy introduces additional concerns: combining datasets can enable inferences about individuals, and removing explicit identifiers does not necessarily prevent reidentification. Security and privacy requirements consequently extend beyond protecting a single database to governing collection, aggregation, analysis, dissemination, and access across the system. (nvlpubs.nist.gov)