A3.4.1 Outline the different types of databases as approaches to storing data. (HL only)

A3.4.1 Outline the different types of databases as approaches to storing data.

Required approaches: NoSQL, cloud, spatial and in-memory databases.

Command term — outline: Give a brief account or summary. Identify the main characteristics of each approach. The explanations and comparisons below develop your understanding; an examination response to “outline” should be more concise.

The big idea

A database stores data so that an application can retrieve, update and manage it. Different approaches suit different data structures, queries and operational needs. A shop might need reliable order records, flexible product descriptions, geographical delivery information and rapidly changing shopping baskets.

No database approach is best for every task. The useful question is: what data must this application store, and what operations must it perform efficiently and reliably?

The four categories in this topic describe different aspects of a database, so they are not mutually exclusive:

ApproachWhat it describesMain question
NoSQLA family of non-relational data models.How is the data organized?
CloudA deployment approach using cloud infrastructure.Where does the database run, and who manages it?
SpatialSupport for spatial data types, indexes and operations.How can locations, shapes and spatial relationships be queried?
In-memoryAn architecture that keeps its primary working data in RAM.How can data be accessed with very low latency?

For example, a cloud-hosted database can use a relational model and support spatial queries. A NoSQL key-value database can also be an in-memory database.


A starting point: what does data look like in a relational database?

A relational database organizes data into tables. Each row represents one record; each column represents an attribute. A schema defines the table structure, including column names, data types and constraints.

Consider a fictional online shop. Its Customers table could contain:

customer_id — integername — textcity — text
101Amina NowakWarsaw
102Leo KowalskiKraków

Its Orders table could contain:

order_id — integercustomer_id — integerorder_date — datetotal_pln — decimal
50011012027-03-04149.90
50021012027-03-0579.00
50031022027-03-05215.50

customer_id is the primary key of Customers: it uniquely identifies each customer. In Orders, it is a foreign key referencing that customer. Customer 101 has two orders, but their name and city need to be stored only once in Customers.

Benefit: separating related entities reduces unnecessary duplication. A foreign-key constraint can prevent an order from referring to a customer who does not exist. A join can combine the tables when an application needs customer names beside order details.

Trade-off: when records have many different attributes, a fixed table structure may require additional tables or schema changes. Relational systems can also support flexible data, such as JSON; the choice is about how naturally the model fits the application.

All examples below are fictional, simplified logical representations of data. They show what a developer might see, rather than the database’s physical arrangement of bytes.


1. NoSQL databases

Overview

NoSQL, often expanded as “Not Only SQL”, refers to a family of non-relational databases. Instead of requiring all data to fit related tables, they may organize it as documents, key-value pairs, wide-column records or graphs.

Many NoSQL systems support flexible schemas or distribution across multiple servers. However, NoSQL does not automatically mean “no structure”, “faster” or “no transactions”. Validation, transaction support and consistency guarantees depend on the particular system and its configuration.

What the data looks like: a document store

A document store could represent two products as the following JSON-like documents:

{
  "product_id": "P101",
  "name": "Waterproof jacket",
  "price_pln": 249.00,
  "sizes": ["S", "M", "L"],
  "waterproof": true
}

{
  "product_id": "P102",
  "name": "Laptop",
  "price_pln": 3299.00,
  "ram_gb": 16,
  "storage_gb": 512
}

Both documents have a product identifier, name and price. Only the jacket has sizes and a waterproof attribute; only the laptop has RAM and storage attributes. Arrays and nested objects can represent related information within one document.

Benefit over one fixed product table: different product categories can have different attributes without adding a column for every possible feature. This can make a changing catalogue easier to extend.

Trade-off: flexible structure still requires rules. If one document uses price_pln and another uses cost, application code may handle them inconsistently. Validation can enforce required fields and types. Embedding duplicated information can also make updates more difficult.

Other NoSQL models

ModelSimplified exampleBenefit and limitation
Key-valuebasket:C101 → {"P101": 2, "P102": 1}A known key retrieves a basket directly. Searching all baskets by their contents may need extra indexes or processing.
Wide-columnPartition: sensor_17; timestamped rows: 10:00 → 18.2°C, 10:01 → 18.4°C.Grouping records around an access pattern can support large volumes of writes and efficient retrieval from a partition. Queries that do not match the partition design may be difficult or expensive.
Graph(Amina)-[FOLLOWS]->(Leo)
(Leo)-[FOLLOWS]->(Sara)
Explicit nodes and relationships suit traversals such as finding friends of friends. This model may add unnecessary complexity to a simple list of independent records.

In the key-value example, product identifiers map to quantities. In the graph example, people are nodes and FOLLOWS connections are directed edges. Wide-column stores organize data around keys and column groupings; they are not simply spreadsheets with many columns.

When is this approach useful?

A document store can suit a varied e-commerce catalogue, a key-value store can suit session lookup, and a graph database can suit social connections. Some NoSQL systems support horizontal scaling, which distributes work across additional servers. Its effectiveness depends on data partitioning and query patterns; relational databases can also scale horizontally.

Sample outline: NoSQL databases use non-relational models such as documents, key-value pairs or graphs. A document database can store products with different attributes, making it suitable for a varied product catalogue.


2. Cloud databases

Overview

A cloud database runs on cloud infrastructure. It can be relational or NoSQL, and it may have spatial or in-memory capabilities. Cloud therefore describes deployment, rather than a separate data format.

A database can be self-managed on a cloud server or supplied as a managed database service, often called Database-as-a-Service (DBaaS). Depending on the service, the provider performs tasks such as patching, backups and failover.

What the data looks like

The Customers and Orders tables shown earlier could be hosted in the cloud without changing their logical structure. The product documents could also be stored in a cloud-hosted document database.

Application → authenticated network connection → cloud database

Orders
order_id | customer_id | order_date | total_pln
5001     | 101         | 2027-03-04 | 149.90

The application communicates with a database service through a network connection. “Cloud” does not mean that the records are publicly accessible; access depends on network and authorization controls.

Benefits compared with running a database on your own server

  • Reduced administration: a managed service can handle routine maintenance, allowing a small development team to spend more time building its application.
  • Flexible capacity: resources can be increased when demand grows and, where supported, reduced later. A shop can prepare for seasonal demand without purchasing hardware for its maximum workload.
  • Availability options: replicas and failover can help a service continue operating when a server fails. These features must be supported and appropriately configured.
  • Lower initial hardware expenditure: an organization can rent capacity rather than buy and maintain its own database server.

Trade-offs

Applications depend on network connectivity and experience network latency. Ongoing charges can become substantial, and changing providers may be difficult. Organizations must still manage permissions, protect credentials, choose suitable data locations and understand backup and recovery arrangements. Managed does not mean maintenance-free, automatically secure or always cheaper.

Typical use: a software-as-a-service (SaaS) application can use a managed cloud database to support customers without operating its own physical database infrastructure.

Sample outline: Cloud databases run on cloud infrastructure and may be provided as managed services. They can offer adjustable capacity and provider-managed backups, making them useful for an online service with changing demand.


3. Spatial databases

Overview

A spatial database supports storing, indexing and querying locations and shapes. Spatial data includes points, such as shop locations; lines, such as roads; and polygons, such as delivery zones. Spatial capabilities can be provided by an extension to a relational database.

What the data looks like

A shop-location table might include a spatial column. The geometry values below are displayed using Well-Known Text (WKT), a readable representation of spatial objects:

shop_idshop_namelocation — spatial value shown as WKT
S01Central branchPOINT(21.0122 52.2297)
S02South branchPOINT(21.0900 52.1600)

For this example, coordinates use WGS 84 and are written as longitude followed by latitude. The coordinate reference system gives the numbers their geographical meaning; the WKT point text alone does not specify it.

A simplified delivery zone could be represented as:

POLYGON((21.00 52.20, 21.10 52.20, 21.10 52.25,
         21.00 52.25, 21.00 52.20))

The repeated first coordinate closes the polygon. A spatial query could ask: “Which shops are within 5 kilometres of this customer?” or “Does this customer’s location fall within the delivery zone?”

Benefits compared with storing coordinates as ordinary numbers or text

  • Built-in spatial operations: the database can calculate distances and test containment or intersection. The application does not need to implement every operation itself.
  • Spatial indexing: an index can narrow down candidate locations or shapes, reducing the number of detailed comparisons needed.
  • Meaningful handling of geometry: points, routes and regions can be processed as spatial objects, rather than as unrelated coordinate strings.

Trade-offs

Spatial indexes require storage and maintenance. Developers must understand coordinate systems, units and geometry validity. For longitude–latitude data, subtracting coordinates does not directly produce a distance in metres: a suitable geographic calculation or projection is needed.

If an application only displays a saved address, ordinary text may be enough. Spatial support becomes valuable when the application must answer geographical questions. Typical uses include GIS, urban planning, environmental monitoring and delivery services.

Sample outline: Spatial databases support geographical objects such as points, lines and polygons. Spatial functions and indexes allow queries such as finding shops within a specified distance of a customer.


4. In-memory databases

Overview

An in-memory database keeps its primary working data in RAM. RAM provides faster access than persistent storage such as an SSD, so this architecture can reduce database latency, especially for workloads involving many small operations.

This does not mean that an in-memory database can never write to disk. It may use transaction logs or snapshots for persistence. Also, conventional disk-oriented databases use RAM caches, so the speed difference depends on the workload, caching, queries and network overhead.

What the data looks like

An in-memory key-value database could hold a shopping session:

Key: session:8fa2
Value:
{
  "customer_id": 101,
  "basket": {"P101": 2},
  "last_page": "/checkout"
}
Expiry: 1800 seconds after the expiry timer is set

The key identifies the session, and the value holds its data. An expiry timer, also called a time to live (TTL), can remove it after a set interval. Expiration is a feature of the chosen system, not a defining property of every in-memory database.

The record’s appearance does not tell you whether it is in RAM or on disk. In-memory describes the storage architecture; the data model could still be key-value, relational or another model.

Benefits compared with relying on disk access for each operation

  • Low latency: keeping working data in RAM reduces storage-access delays, which can help rapidly changing leaderboards, sessions and live analytics.
  • High throughput: suitable workloads can process many operations per second when disk access would otherwise be a bottleneck.
  • Reduced load on another database: an in-memory cache can hold frequently requested results, reducing repeated queries to the main database.

Trade-offs

RAM is volatile: its contents are lost when power is removed. Recovery therefore depends on persistence and replication arrangements. Snapshots may omit changes made since the last snapshot; logging policies affect both performance and the amount of recent data at risk.

RAM also generally costs more per unit of capacity than persistent storage. Large archives may be more economical on disk. If the in-memory system acts as a cache, the application must prevent or tolerate stale data: cached values that no longer match the authoritative record.

Example decision: an application might keep recoverable session data in memory while recording completed orders durably in a relational database. An in-memory database can also store important persistent data when its durability and recovery design meets the application’s requirements.

Sample outline: In-memory databases keep their primary working data in RAM to provide low-latency access. They suit rapidly changing data such as live leaderboards, but require persistence mechanisms to recover data after power loss.


Comparing the approaches

Choose an approach by linking a requirement to a technical feature and its consequence. These are comparisons against relevant alternatives, rather than a ranking of four exclusive choices.

RequirementUseful approachWhy it helpsTrade-off to consider
Maintain related customer and order records.Relational database.Keys, constraints and transactions help preserve valid relationships and consistent updates.Schema design and joins require planning.
Store products with widely varying attributes.NoSQL document database.Documents can represent different attributes and nested data naturally.Validation and duplicated data still require management.
Operate an online service with limited infrastructure staff.Managed cloud database.The provider can handle routine infrastructure tasks and offer capacity adjustments.Network dependence, recurring costs and provider dependence.
Find nearby shops and check delivery zones.Spatial database capabilities.Spatial types, functions and indexes support geographical queries.Coordinate systems and spatial indexes add complexity.
Update and retrieve live scores very frequently.In-memory database.RAM access can reduce latency for frequent operations.Memory cost, capacity and recovery requirements.

An online shop could combine these approaches: relational tables for orders, documents for product details, spatial capabilities for deliveries and in-memory storage for sessions. Any of these could run in the cloud. Combining databases also adds integration and consistency work, so a small application may be better served by one database with suitable capabilities.

Using the command term accurately

Outline requires a brief account of the main characteristics. It does not require the full comparison developed in this article. Use the question’s wording and mark allocation to judge how much detail to include.

Less effective: “In-memory databases are fast and good for games.” This is vague: it does not identify the storage medium or the relevant kind of workload.

More effective outline: “In-memory databases keep their primary working data in RAM, providing low-latency access. They are suitable for frequently updated game leaderboards.” This identifies the defining feature and connects it to a relevant use.

Additional explanatory depth: “RAM access avoids many delays associated with retrieving data from persistent storage, which can reduce response time for frequent score updates. However, RAM is volatile, so recovery requires persistence or another recoverable copy.” This explains the mechanism and a limitation; it develops understanding beyond a brief outline.

When explaining a benefit, make the reasoning explicit: requirement → database feature → practical consequence. For example: products have different attributes → documents support varying fields → new product categories can be added without adding a relational column for every possible attribute.