6:45 a.m. The purchasing report is late again.
Imagine a restaurant group with 12 locations. Buyers need the next day's ingredient forecast before 7 a.m., but the report now arrives after 8. Someone proposes a data lake, a streaming platform, and a distributed compute cluster.
The CTO traces the slow job first. Every morning it rereads three years of order history from unpartitioned files, even though only yesterday's orders changed. This illustrative business has a real performance problem. It does not yet have proof that it needs a data lake, or any larger big data architecture. This guide shows how to get that proof before anyone builds one.
Quick answer: Build a data lake or a larger big data architecture only when measured workload pressure exceeds a simpler platform's limits for processing time, concurrency, resilience, data variety, retention, or required latency. Choose the smallest architecture that meets a named service level, test it with representative loads, and include ownership, observability, security, and cost from the first design.
Big data architecture starts with pressure, not volume
There is no universal row count that makes a system “big.” A dataset can be modest and still require distributed processing because the latency target is seconds, the sources arrive continuously, or hundreds of users query it together. A much larger archive may work well with daily batch processing and inexpensive object storage.
A big data architecture decision therefore needs two measurements:
- The workload pressure today and at a credible future peak.
- The service level the business actually needs.
AWS makes the same point in its Data Analytics Lens: there are many valid data platform methods, and the chosen approach must fit the organization's principles, design objectives, resources, and experience. Big data architecture is a tradeoff, not a maturity badge.
For the restaurant group, the service level is concrete: the forecast must be ready before buyers place orders at 7 a.m. "Make the platform modern" is not a service level. "Process yesterday's branch and delivery orders by 6:45 a.m., with a recoverable failure path" is.
The seven pressure tests for big data architecture
Run these big data architecture tests before choosing products. Record evidence, not adjectives.
- Processing pressure: Does the current job miss its required completion window? Capture runtime by stage, input size, and the measured bottleneck.
- Latency pressure: Must the business react in seconds or minutes? Record the decision deadline and the real delay from event to action.
- Concurrency pressure: Do users or jobs compete for resources? Measure peak queries, queue time, throttling, and failed work.
- Variety pressure: Do formats and schemas change faster than the current model can absorb? Count source types, schema changes, and parse failures.
- Reliability pressure: Could one machine or service failure breach the service level? Define the recovery objective, retry behavior, and operational impact.
- Retention pressure: Must raw and historical data remain queryable at a reasonable cost? Inspect retention rules, access frequency, and storage growth.
- Isolation pressure: Do teams or workloads need separate security, cost, or compute boundaries? Document access models, workload ownership, and cost attribution.
Big data architecture is worth evaluating when several tests show sustained pressure and the value of meeting the service level exceeds the cost of added complexity. A single slow query is a tuning problem until evidence shows otherwise.
The restaurant case fails one pressure test at first: processing time. Its latency target is measured in hours, not seconds; its data is mostly structured; and only a small operations team uses the result. That evidence points to incremental loading, partitioning, and query repair before a distributed stack.
The architecture rule: Scale the design only when a measured constraint survives reasonable tuning and the business value of removing it exceeds the operating cost you will inherit.
Test the processing window
Measure each stage of the pipeline. Data extraction, transfer, transformation, joins, model execution, storage writes, and dashboard refresh may have different constraints. A big data architecture will not repair an unfiltered source query or an unnecessary full reload by itself.
Before expanding big data architecture, start with file format, partitioning, indexing, incremental processing, and query design. AWS recommends choosing formats around access patterns, using compression, and partitioning data so engines can avoid unnecessary reads. If those changes still cannot meet the processing window under representative load, distributed compute becomes an evidence-based option.
In the restaurant story, changing the job to process only new orders may return the forecast to the 6:45 a.m. window. If it does, the team has solved the business problem without inheriting a cluster. If it still fails under the expected branch and order growth, the same measurement becomes evidence for the next architecture stage.
Separate business latency from technical excitement
Ask when the next action becomes less valuable. Fraud intervention may need a rapid signal. A board pack may not. Big data architecture should reflect that difference.
Streaming big data architecture adds event ordering, late data, replay, state, back pressure, schema evolution, and continuous operations. Apache Kafka provides APIs for producing, consuming, and processing event streams, including stateful operations, joins, windows, and event time. Those capabilities solve real problems, but the team inherits their operating questions too.
If an hourly micro-batch meets the decision deadline, do not create a streaming big data architecture merely because the source emits events.
Restaurant orders arrive continuously, but tomorrow's purchasing decision happens once each morning. Continuous source data does not automatically create a real-time business need. If the group later adds live fraud checks or minute-by-minute courier allocation, that new decision may justify a different path.
Measure concurrency and isolation
In a big data architecture, one workload may be fast alone and unreliable when analysts, models, and reporting jobs run together. Capture peak demand, queue delay, resource contention, and failure patterns.
This is where a big data architecture that separates storage from compute can help. Independent compute pools can serve transformations, exploration, and production reporting without forcing every workload into one capacity decision. The gain is isolation and elasticity. The cost is more configuration, monitoring, policy, and financial control.
Choose a stage, not a fashionable stack
Big data architecture can grow through stages. The labels below describe patterns, not mandatory products.
Stage 1: One managed analytical engine. Use a relational database or warehouse with scheduled transformations when structured data and batch decisions fit comfortably. The main danger is business logic spreading across individual reports.
Stage 2: Durable storage with a separate query layer. This is where a data lake usually enters. Add object storage with a warehouse or lakehouse query layer when raw history, changing formats, or economical retention matters, which means the retention and variety tests show sustained pressure. The main danger is creating two stores without clear ownership and governance.
Stage 3: Distributed batch processing. Introduce distributed compute when representative jobs still exceed the practical processing window after sound tuning. The main danger is giving the team a cluster it lacks the capacity to operate.
Stage 4: Continuous event processing. Add event streaming and stream processing when the value of an action decays before a batch can finish. The main danger is permanent operational complexity for a problem that did not require it.
Apache Spark's cluster overview shows what distributed execution introduces: a driver coordinates applications, a cluster manager allocates resources, and executors run tasks and store data. A big data architecture at this stage needs capacity planning, retry behavior, logs, deployment controls, and people who can diagnose each layer.
The right stage is the lowest one that passes the pressure tests with acceptable headroom.
That stage can change. A 12-location group may fit a managed warehouse with incremental jobs. A future operation with hundreds of locations, several delivery partners, live stock updates, and independent data science workloads may earn separate storage and compute. The architecture grows because the pressure changed, not because the original diagram looked ambitious.
What proof of need looks like
Proof is a short written case, not a vendor demo. Before anyone approves a data lake or the next architecture stage, ask the team for one page that answers six questions:
- The decision. Which business decision is late, wrong, or impossible today, and who owns it?
- The service level. What deadline, freshness, or recovery target must the platform meet?
- The baseline. How does the current system perform against that target under representative load?
- The tuning already tried. Which fixes, such as incremental loads, partitioning, or query repair, were tested, and what did they change?
- The failing test. Which of the seven pressure tests still fails, and by how much?
- The cost of the answer. What will the new stage cost to run, including the people who operate it, and who owns that budget?
If the team cannot name the failing test, the proposal is not ready. For the restaurant group, the page shows one decision (the 7 a.m. purchasing order), one failing test (processing time), and one untested fix (incremental loading). That page says: fix the job first, then measure again.
A reference shape for big data architecture
Whatever products you choose, keep six responsibilities visible:
- Sources and contracts. Define ownership, schema, arrival expectations, and permitted use.
- Ingestion. Support retries, idempotency, quarantine, and traceable failure.
- Storage. Separate raw, validated, and consumption ready data where the distinction serves recovery and trust.
- Compute. Match batch, interactive, or streaming execution to the decision deadline.
- Serving. Give BI, applications, analysts, and AI products stable interfaces and definitions.
- Control. Apply access, lineage, quality, observability, retention, and cost attribution across the big data architecture.
This big data architecture shape prevents a common mistake: drawing storage and compute boxes while leaving ownership, recovery, and consumption implicit.
Cost belongs inside the big data architecture
Cloud elasticity changes how cost appears. It does not remove cost. A credible big data architecture model includes:
- Storage by class, retention, replication, and region.
- Compute by workload, schedule, peak, and idle time.
- Data transfer between services, regions, and consumers.
- Orchestration, catalog, observability, security, and support.
- Engineering time for upgrades, incidents, performance work, and data recovery.
AWS's analytics cost optimization guidance recommends measuring cost at each processing step or pipeline branch. It also recommends matching compute and storage to usage patterns and removing or downsizing unused jobs and resources.
That makes cost a feedback loop. Tag each workload to an owner, report unit cost such as cost per pipeline run or report refresh, and investigate change. Big data architecture economics become manageable when teams can see which decision consumes the resources.
When big data architecture is not the next move
Stay with a simpler design when the current system meets the service level after reasonable tuning, batch timing fits the decision, most data is structured, concurrency is limited, and the team cannot operate distributed systems reliably.
In that case, improve the data model, incremental loads, tests, definitions, and workload scheduling. Revisit big data architecture when the pressure tests change. Simplicity is not technical debt when it is measured, documented, and reversible.
For the late restaurant report, the first win is not a lake. It is a purchasing forecast that arrives before the buyer needs it, with evidence showing why the simpler design still works. That is what a sound architecture decision looks like from the business side.
Frequently asked questions about big data architecture
What does a big data architecture include?
A complete big data architecture connects sources, ingestion, storage, processing, serving, and control. The control layer includes security, quality, lineage, observability, retention, recovery, and cost ownership. The products can vary. What matters is that each responsibility has a clear boundary, service level, and owner.
Do we need a data lake or a data warehouse?
Start with the decision, not the store. A data warehouse suits structured data and repeatable reporting, which covers many small and mid-sized businesses. A data lake earns its place when you must keep large volumes of raw or semi-structured data at low cost, such as logs, events, files, or history for data science, and query it later in ways you cannot fully predict. Many teams now combine the two in a lakehouse. Whichever you choose, the retention and variety pressure tests should show sustained pressure first.
When does a business actually need big data architecture?
A business needs big data architecture when measured processing, latency, concurrency, reliability, variety, retention, or isolation demands exceed what a simpler design can meet after reasonable tuning. Data volume alone is not enough. The decision should be tied to a business service level and tested under representative peak conditions.
Does big data architecture always require real-time streaming?
No. Many big data architecture workloads are best served by scheduled or incremental batch processing. Streaming is justified when the value of action falls materially before a batch can finish. If an hourly update meets the decision deadline, continuous event processing may add cost and operating risk without adding business value.
How can a team control big data architecture costs?
Separate workloads, assign each one an owner, measure unit costs, match storage and compute to actual access patterns, and remove idle or duplicated processing. A big data architecture should expose cost by pipeline, service, or consumer. Cost visibility turns architecture economics into an operating decision instead of a surprise on the monthly bill.
For the operating controls around the platform, read Xyric's data governance starter model. The Optimize track covers data strategy, engineering, analytics, and AI for existing operations.
Book a discovery call to run one workload through the seven big data architecture pressure tests before committing to a data lake or any new platform.
Sources
- AWS Well Architected Data Analytics Lens, Amazon Web Services, accessed 2026-09-08.
- Modern data architecture, Amazon Web Services, accessed 2026-09-08.
- Cost optimization for data analytics, Amazon Web Services, accessed 2026-09-08.
- Apache Spark cluster mode overview, Apache Spark, accessed 2026-09-08.
- Apache Kafka introduction and APIs, Apache Kafka, accessed 2026-09-08.


