An insurer hands you a policy database, a folder of adjuster claim notes, and a shelf of accident-scene photos. Only one of these drops straight into a regression. Knowing which, and why, is the whole point of this outcome.
What "structured" really means. Structured data lives in a table. Think of a spreadsheet or a relational database. Every column has a name, a defined data type, and a consistent meaning. Every row is a single record. You can query it, sort it, and feed it to a model without reshaping it.
The defining trait is a predefined schema. Before any record arrives, you already know the fields: PolicyID, DriverAge (integer), Region (category), AnnualPremium (numeric). Each new record slots into that fixed structure.
KEY: If you can describe the data as "rows are observations, columns are variables, and every value in a column shares one data type," it is structured.
What "unstructured" means. Unstructured data has no row-and-column organization and no fixed schema. The content does not decompose into named fields on its own.
Common mistakes
- Calling any messy data "unstructured." Data can be structured yet dirty (missing values, typos). Structure is about the schema (rows and columns), not cleanliness. A table with missing entries is still structured.
- Treating semi-structured as unstructured. JSON, XML, and log files have keys or tags. They are semi-structured and parse into tables; do not lump them with free text or images.
- Assuming unstructured data cannot be modeled. It can, but only after you extract features to build structured variables. The narrative "no injuries reported" becomes a usable indicator flag.
Bottom line
- Structured data fits a fixed row-and-column table with a predefined schema and one data type per field.
- Unstructured data has no predefined tabular format: free text, images, audio, video, sensor streams.
- Semi-structured data carries tags or markers (JSON, XML, HTML, log files) but not rigid rows and columns.
- In structured data, each row is one observation (record) and each column is one variable (field).
Exam shortcut
Classify by schema: fixed named columns is structured, tags or keys with variable shape is semi-structured, free-form content with no fields is unstructured. When a prompt lists unstructured sources (text, images, audio, video), your answer almost always includes "extract features to create structured variables first," since standard models need tabular input.
The full lesson (about 1,587 words, 11 min read) adds 2 worked examples, all 6 common mistakes, a self-check, free in the app.
Learning objectives
- 2a
Browse all free Exam PA lessons or jump into free Exam PA practice questions.