Learning Module 01 · Week 1 · Individual mission
Diagnose Big Data Before Choosing a Tool
Decide whether a finance problem needs big-data methods before reaching for a fashionable tool.
- 1BriefingCurrent
- 2Worked exampleNext
- 3Case boardNext
- 4CheckpointsNext
- 5ReflectionNext
- 6CompleteNext
Start here
Briefing
Big data is not defined by file size alone. In finance, the combination of scale, speed, form, reliability, and decision value matters.
Use the provisional 5V core as a diagnostic. Treat additional Vs as broader teaching frameworks, not as an approved universal taxonomy. The V set itself grew from three to five to seven or eight, so treat broader Vs as a timeline of teaching extensions.
Your answers to each question are kept only while this page is open and reset when you refresh. Your completed-mission progress and Decision XP are saved in this browser on this device only, never sent to a server. Clearing your browser data, or using Clear my progress, removes them.
Game economy
Build Evidence Momentum without risking Decision XP
You begin with 3 Evidence Momentum. Any held points return when you solve that checkpoint. Evidence Momentum never falls below zero and never changes assessment marks. Complete the mission to open one transparent reward draw.
What you will practise
- Distinguish a merely large dataset from a big-data problem using a provisional 5V core.LO1
- Explain why broader V frameworks are context-dependent rather than a fixed official list.LO2
7 terms for this mission
- Volume
- The scale of data to be stored or processed.
- Velocity
- The speed at which data arrives and must be acted on.
- Variety
- The mix of structured, semi-structured, and unstructured data.
- Veracity
- The reliability, quality, and uncertainty of the data.
- Value
- The useful decision or outcome the data can support.
- Variability
- How much the meaning or flow of data shifts over time, one of the extension Vs added beyond the core five.
- Visualization
- How clearly data can be presented so a decision-maker can interpret and act on it, another extension V beyond the core five.
Every organisation, person, figure, and dataset in this mission's scenario is invented. Real institutions, laws, and standards are named only as general context.
Sources and further reading
Every organisation, person, figure, and dataset in this mission's scenario is invented. Real institutions, laws, and standards named in this mission, including in the sources below, appear only as general context, and the sources support the concepts and methods, not the events in the scenario.
Walker, T., Davis, F., & Schwartz, T. (Eds.) (2022). Big Data in Finance: Opportunities and Challenges of Financial Digitalization (1st ed.). Palgrave Macmillan. doi:10.1007/978-3-031-12240-8
A finance-specific overview of what makes financial data big, covering scale, data form, data quality and decision value rather than tool choice alone.
Chang, W. L., & Grady, N. (2019). NIST Big Data Interoperability Framework: Volume 1, Definitions (NIST Special Publication 1500-1r2). National Institute of Standards and Technology. doi:10.6028/NIST.SP.1500-1r2
A standards body's consensus definition, treating volume, velocity, variety and variability as the core set and veracity and value as further terms, which shows why the V list is a working framework rather than one official list.
Ngo, V. M., Huynh, T. L. D., Nguyen, P. V., & Nguyen, H. H. (2022). Public sentiment towards economic sanctions in the Russia-Ukraine war. Scottish Journal of Political Economy, 69(5), 564-573. doi:10.1111/sjpe.12331
A published example of building a machine-learning sentiment measure from close to a million unstructured social-media posts across 108 countries, showing what volume and variety look like once text replaces a tidy table.
Soldatos, J., & Kyriazis, D. (Eds.) (2022). Big Data and Artificial Intelligence in Digital Finance: Increasing Personalization and Trust in Digital Finance using Big Data and AI. Springer. Open access. doi:10.1007/978-3-030-94590-9
Describes how institutions handle high-volume, high-velocity and mixed-format data, useful for judging when ordinary processing is no longer sufficient.