High School AdvancedModule A5Lesson 2 of 10Evidence, Provenance, Health, and Privacy

A5.2 Data Sources for Detection

Learn how fictional identity, endpoint, network, DNS, email, application, cloud, supplier, administrative, support, and source-health data contribute different evidence—and why provenance, field meaning, freshness, completeness, timing, coverage, transformation, duplication, privacy, ownership, and failure behavior determine what a detection can responsibly claim.

Lesson Progress

Data Sources for Detection

High School AdvancedA5: Detection Engineering • Lesson 2 of 10

20% complete

Readiness Check

Before You Start

0/6 ready

Professional Hook

A Green Data Source Can Still Produce Weak Evidence

A fictional network collector is connected and reports a Green heartbeat. After a parser update, event volume drops, one field disappears, and application correlation becomes delayed. The detection platform is still receiving records—but the source may no longer answer the same defender questions with the same confidence.

Weak conclusion

“The collector is Green, so detections using this source are healthy.”

Strong conclusion

“The fictional source is connected, but freshness, completeness, schema, field mapping, volume, transformation, and application-correlation health require validation before normal confidence is restored.”

Data presence is not the same as evidence quality, and evidence quality is not the same as a confirmed conclusion.

Exactly Five Learning Objectives

What You Will Be Able to Do

Objective 1

Classify fictional identity, endpoint, network, DNS, email, application, cloud, supplier, administrative, support, and source-health data according to the defender questions they can support.

Objective 2

Evaluate fictional data sources using provenance, field meaning, freshness, completeness, timing, schema, transformation, duplication, coverage, privacy, ownership, retention, and failure behavior.

Objective 3

Distinguish source availability from source quality and explain how delayed, missing, transformed, duplicated, or inconsistent evidence changes detection confidence.

Objective 4

Design a fictional detection data-source catalog, source-health model, evidence-provenance map, privacy plan, field dictionary, and coverage-gap register.

Objective 5

Create a portfolio-ready fictional source-governance package that connects mission risks, defender questions, evidence sources, health requirements, limitations, owners, and review triggers.

Why This Matters

Detections Can Only Be as Responsible as Their Evidence

Fictional detection logic may be perfectly documented and still fail if its sources arrive late, omit required identities, change field meaning, duplicate records, lose application context, use stale enrichment, overcollect personal information, or fail silently during recovery. Data-source engineering gives defenders the evidence foundation needed to interpret every later alert.

Question-source fit

Choose fictional evidence according to the defender question rather than source popularity or volume.

Health-aware confidence

Change fictional confidence and analyst guidance when required sources are delayed, incomplete, conflicting, or missing.

Privacy-aware usefulness

Collect only fictional fields needed for approved detection and triage decisions.

Core Framework

The S-O-U-R-C-E Method

S — Start with the defender question

Define which fictional decision needs evidence and what no source can prove alone.

O — Observe provenance and ownership

Identify the fictional producing system, owner, schema, collection, transformation, enrichment, storage, and access path.

U — Understand fields and timing

Document fictional field meaning, direct or derived status, event time, collection time, processing time, and clock limits.

R — Review health and coverage

Measure fictional connectivity, freshness, completeness, queue, schema, duplication, populations, environments, and blind periods.

C — Control privacy and change

Use fictional purpose-based fields, access, retention, deletion, source-change notifications, versioning, and review triggers.

E — Explain evidence limits

State which fictional conclusions the source supports, which alternatives remain, and how degraded conditions affect confidence.

Decision-ready source statement

This fictional source supports a documented defender question through defined fields, provenance, timing, coverage, health, privacy, ownership, retention, transformation, and limitations. When source conditions degrade, detection confidence and analyst guidance change according to an approved plan.

Advanced Vocabulary

Terms for Detection Data Sources

Detection data source

A fictional identity, endpoint, network, DNS, email, application, cloud, supplier, administrative, support, or health record used to answer a defender question.

Provenance

Fictional information describing where evidence originated, which system or role produced it, how it moved, and which transformations occurred.

Field dictionary

A fictional record explaining field names, meanings, allowed values, timing, ownership, privacy, dependencies, and known limitations.

Freshness

The fictional amount of time between an event or state change and when the evidence becomes available for detection or review.

Completeness

The fictional degree to which expected records and required fields are present for the relevant scope and period.

Event time

The fictional time associated with when an activity or state change occurred.

Collection time

The fictional time associated with when a source or collector received the evidence.

Processing time

The fictional time associated with when evidence was normalized, enriched, transformed, stored, or made available.

Clock alignment

The fictional degree to which evidence sources use sufficiently consistent time references for reliable sequencing and correlation.

Schema

A fictional structure describing which fields, types, values, relationships, and constraints a source provides.

Normalization

A fictional transformation that maps differently structured records into a common model for comparison and correlation.

Enrichment source

A fictional source that adds context such as asset value, identity role, service owner, device class, location concept, change state, or source health.

Transformation

A fictional operation that parses, maps, aggregates, filters, masks, groups, or derives fields from original evidence.

Duplication

A fictional condition in which the same underlying activity appears in more than one record, source, collector, or transformed dataset.

Coverage

The fictional extent to which evidence represents the relevant identities, devices, services, environments, zones, time periods, states, and workflows.

Blind spot

A fictional area where relevant evidence is unavailable, unreliable, delayed, inaccessible, or outside the designed scope.

Blind period

A fictional time window during which a source was missing, delayed, unhealthy, or unable to support normal confidence.

Source health

Fictional evidence about connectivity, event freshness, completeness, volume, queue age, clock, schema, transformation, duplication, storage, access, and blind periods.

Source owner

The fictional role accountable for field meaning, health, change communication, access, privacy, retention, support, and lifecycle.

Authoritative source concept

A fictional source designated as the approved record for a specific state or decision, while still requiring health and scope validation.

Corroborating source

A fictional source that supports, challenges, or adds context to an observation from another source.

Derived field

A fictional value calculated from one or more original fields rather than recorded directly by the producing system.

Retention

The fictional period for which evidence is stored according to purpose, privacy, access, operational, legal, and deletion requirements.

Source review trigger

A fictional event requiring revalidation, such as schema, platform, application, identity, network, supplier, privacy, retention, ownership, or mission change.

Instructional Section 1

Compare Eleven Detection Source Categories

Identity evidence

Defender questions

Which fictional user, service, supplier, privileged, emergency, or recovery identity authenticated, received authority, changed roles, or lost access?

Strong fictional fields

Identity category, role, group, approval, authentication result, authorization result, session, expiration, revocation, owner, and source health.

Strengths

Supports identity lifecycle, privilege, assignment, session, revocation, and policy-related questions.

Limitations

Successful authentication does not prove the action, destination, object, purpose, or outcome was authorized.

Privacy boundary

Prefer role, identity category, owner group, and decision context over unrelated personal profile details.

Health requirements

Track assignment freshness, group synchronization, revocation delay, clock, missing records, and source changes.

Endpoint evidence

Defender questions

Which fictional managed device, service device, administrative device, or application process produced a local state or behavior?

Strong fictional fields

Device identity, owner, class, health, process category, application, policy state, event category, result, time, and source health.

Strengths

Can add local device, process, configuration, and application context unavailable from network evidence alone.

Limitations

A local event does not automatically prove network reachability, user intent, service authorization, or wider impact.

Privacy boundary

Collect only fields necessary for the approved defender question and avoid unrelated user activity detail.

Health requirements

Track agent availability conceptually, last event, version, queue, clock, schema, duplicate reporting, and blind periods.

Network evidence

Defender questions

Which fictional source, destination, direction, service category, policy result, timing, duration, or volume relationship occurred?

Strong fictional fields

Source group, destination group, identity context, service category, direction, policy result, timing, duration, volume class, sensor, and source health.

Strengths

Supports communication-path, segmentation, firewall, remote-access, wireless, and service-relationship questions.

Limitations

Network evidence may not prove application action, object authorization, user intent, content meaning, or business outcome.

Privacy boundary

Use grouped identities, service categories, and minimized destination context when exact detail is unnecessary.

Health requirements

Track sensor coverage, freshness, queue age, clock, policy version, packet or event loss concept, transformation, and encrypted boundaries.

DNS evidence

Defender questions

Which fictional requester group asked which question category, through which resolver, and received which response category under which policy and cache state?

Strong fictional fields

Requester group, resolver, question category, response category, cache state, policy result, timing, source health, and application correlation.

Strengths

Supports naming, service-discovery, resolver-policy, unexpected-resolution, cache, and service-dependency questions.

Limitations

Successful resolution does not prove destination authorization, service health, application safety, or user intent.

Privacy boundary

Exact naming activity may reveal sensitive interests or internal architecture; use purpose-based minimization.

Health requirements

Track resolver availability, event freshness, policy version, cache context, clock, queue age, authoritative relationship, and blind periods.

Email and messaging evidence

Defender questions

Which fictional message, sender category, recipient group, routing service, policy decision, supplier relationship, attachment category, or delivery result occurred?

Strong fictional fields

Sender category, recipient group, service identity, route category, policy result, delivery state, attachment class, supplier, time, and source health.

Strengths

Supports delivery, policy, supplier, routing, and message-workflow questions.

Limitations

Message metadata does not automatically prove content meaning, user intent, delivery reading, or downstream action.

Privacy boundary

Prefer categories, policy outcomes, and routing context over unnecessary message content or personal details.

Health requirements

Track source availability, queue age, delivery delay, duplicate events, supplier status, schema, and retention.

Application evidence

Defender questions

Which fictional user or service performed which approved operation on which object or workflow, with which result and business state?

Strong fictional fields

Identity, role, assignment, object category, operation, old state, new state, result, application version, change, owner, and source health.

Strengths

Often provides the strongest context for service authorization, workflow state, object scope, and business outcome.

Limitations

Application logs may omit network path, device health, source provenance, or external supplier state.

Privacy boundary

Minimize object detail and personal content while preserving the decision-relevant operation and result.

Health requirements

Track event freshness, required fields, transaction coverage, queue, schema, application version, clock, and blind periods.

Cloud-service evidence

Defender questions

Which fictional cloud identity, service, resource category, administrative action, policy decision, configuration change, or access relationship occurred?

Strong fictional fields

Identity category, service, resource class, action category, policy result, role, change, region concept, time, source health, and owner.

Strengths

Supports service, identity, administrative, policy, configuration, and provider-context questions.

Limitations

Provider records may reflect only one layer and may not prove application purpose, user intent, or downstream impact.

Privacy boundary

Use resource and identity categories where exact names are unnecessary.

Health requirements

Track export or collection delay, schema changes, provider status, duplicate records, clock, scope, permissions, and retention.

Supplier evidence

Defender questions

Which fictional supplier identity, request, result, support action, remote session, delivery state, or contractual responsibility is relevant?

Strong fictional fields

Supplier identity, sponsor, request category, result category, support ticket, destination, session, contract state, time, source health, and owner.

Strengths

Supports shared-responsibility, external access, request-result correlation, support, and continuity questions.

Limitations

Supplier-provided evidence may not prove internal authorization, full processing, independent operation, or user impact.

Privacy boundary

Limit data to the approved supplier relationship, service need, owner, and defensive purpose.

Health requirements

Track supplier availability, delay, schema, support status, correlation, coverage, contract changes, and blind periods.

Administrative and change evidence

Defender questions

Which fictional privileged identity approved or performed which change, on which service or control, with which validation and rollback?

Strong fictional fields

Administrator identity, device, role, approval, change identifier, target category, old state, new state, result, rollback, time, and source health.

Strengths

Supports privileged action, change, maintenance, emergency, and recovery context.

Limitations

A change record does not prove implementation success, complete scope, or correct business outcome.

Privacy boundary

Preserve accountability without exposing unnecessary internal configuration or personal details.

Health requirements

Track approval freshness, role state, change-system availability, session evidence, clock, closure, and retrospective review.

Support and user-confirmation evidence

Defender questions

Which fictional user or owner reported impact, confirmed a change, requested support, accepted a resolution, or disputed an outcome?

Strong fictional fields

User group, service, issue category, assignment, ticket, confirmation, impact, accessibility need, time, owner, and source health.

Strengths

Adds mission, user-impact, support, accessibility, and business-outcome context.

Limitations

A report or confirmation may be incomplete, delayed, mistaken, or limited to one user's experience.

Privacy boundary

Use issue and impact categories instead of unnecessary personal or content detail.

Health requirements

Track ticket freshness, duplicate reports, assignment, closure, communication delay, sampling limits, and source availability.

Source-health evidence

Defender questions

Can the fictional source support normal confidence for the relevant scope, fields, time, and decision?

Strong fictional fields

Connectivity, last event, event volume, queue age, clock, schema, transformation, duplication, storage, access, blind period, and owner.

Strengths

Explains whether other evidence is timely, complete, interpretable, and available.

Limitations

Healthy source metrics do not prove the underlying behavior is safe or harmful.

Privacy boundary

Usually requires limited operational metadata rather than personal activity detail.

Health requirements

Source-health evidence also needs ownership, freshness, completeness, and independent validation.

Instructional Section 2

Trace the Ten-Stage Evidence Provenance Chain

1. Event or state

Review question

What fictional activity, decision, request, result, configuration, identity state, or service condition occurred?

Fictional evidence

Original producing system context and event-time concept.

Risk

The source may not record every relevant action or may represent only one layer.

2. Source generation

Review question

Which fictional system or role created the evidence and according to which schema and version?

Fictional evidence

Source identifier, owner, schema, version, event category, required fields, and generation conditions.

Risk

Different versions may use different fields or meanings.

3. Collection

Review question

When and how was the fictional evidence received by a collector or export process?

Fictional evidence

Collection time, collector identity, channel concept, queue, connectivity, and missing-period record.

Risk

Collection may be delayed, partial, duplicated, or interrupted.

4. Parsing and normalization

Review question

How was the fictional record interpreted and mapped into common fields?

Fictional evidence

Parser version, field mapping, schema result, parse status, dropped fields, and transformation notes.

Risk

Field meaning may change or data may be incorrectly mapped.

5. Enrichment

Review question

Which fictional identity, asset, service, owner, change, peer, or risk context was added?

Fictional evidence

Enrichment source, version, join key concept, freshness, confidence, and owner.

Risk

Stale or incorrect enrichment can create misleading context.

6. Storage

Review question

Where and for how long is fictional evidence retained, and who may access it?

Fictional evidence

Storage category, retention, access roles, integrity checks, privacy, and deletion.

Risk

Retention or access may exceed purpose, while unavailable storage creates blind periods.

7. Detection evaluation

Review question

Which fictional logic version evaluated which fields and source-health state?

Fictional evidence

Detection version, required fields, evaluation time, source state, result, and confidence.

Risk

The logic may run with missing or degraded context without clearly indicating the limitation.

8. Alert presentation

Review question

Which fictional evidence, enrichment, severity, confidence, limits, and next questions reached the analyst?

Fictional evidence

Alert version, displayed fields, source health, explanation, owner, and triage guidance.

Risk

Important limitations may be hidden while irrelevant detail overwhelms the analyst.

9. Analyst decision

Review question

Which fictional evidence was reviewed and which conclusion or action was recorded?

Fictional evidence

Evidence request, alternatives, confidence, scope, impact, decision, escalation, and closure.

Risk

The analyst may treat alert text as fact or miss source-health limitations.

10. Feedback and lifecycle

Review question

How did the fictional result improve source, logic, testing, tuning, documentation, ownership, or retirement?

Fictional evidence

Outcome label, defect, change, test, metric, owner, due date, validation, and review trigger.

Risk

Without feedback, repeated false positives, false negatives, and source gaps remain unresolved.

Instructional Section 3

Separate Event, Collection, and Processing Time

Time conceptFictional meaningDefender useImportant limitation
Event timeWhen the fictional activity or state change occurred according to the producing source.Sequence identity, service, network, DNS, application, supplier, and change behavior.The source clock may be wrong, delayed, or based on a different state transition.
Collection timeWhen a fictional collector or export process received the record.Measure transport delay, backlog, and blind periods.Collection time does not tell when the underlying activity occurred.
Processing timeWhen fictional parsing, normalization, enrichment, storage, or detection evaluation occurred.Measure pipeline delay and identify transformation bottlenecks.Processing time can be much later than event time.
Alert timeWhen the fictional detection result became visible to an analyst.Measure end-to-end detection delay and response opportunity.Alert time alone does not reveal where delay occurred.
Owner-confirmation timeWhen a fictional service, identity, supplier, or user owner confirmed context or outcome.Support authorization, impact, false-positive, false-negative, and closure review.Confirmation may be delayed, incomplete, or limited to one perspective.
Recovery timeWhen fictional evidence, service, policy, or source health returned to an approved state.Define blind periods, reprocessing, reassessment, and closure.A source returning does not prove historical gaps are repaired.

Instructional Section 4

Evaluate Ten Source-Health Dimensions

Connectivity

Review question

Can the fictional collector, export, API, agent, or pipeline currently communicate?

Strong fictional evidence

Connection state, last successful exchange, failure reason, owner, and independent confirmation.

Important limitation

Connectivity does not prove current events or complete fields.

Freshness

Review question

How old is the newest fictional event or state relevant to the defender question?

Strong fictional evidence

Last-event time, expected delay, actual delay, queue age, source type, and state.

Important limitation

A fresh event does not prove complete coverage.

Completeness

Review question

Are expected fictional records and required fields present for the relevant scope and period?

Strong fictional evidence

Expected volume range, field presence, missing categories, comparison sources, and blind-period record.

Important limitation

Normal volume does not prove every critical event is present.

Clock alignment

Review question

Can fictional events from different sources be sequenced with sufficient confidence?

Strong fictional evidence

Clock status, known offset, event time, collection time, processing time, and uncertainty.

Important limitation

Aligned clocks do not prove semantic correlation.

Schema stability

Review question

Did fictional field names, types, values, relationships, or required fields change?

Strong fictional evidence

Schema version, change notice, parser status, field dictionary, test results, and owner review.

Important limitation

A valid schema does not prove correct field meaning.

Transformation quality

Review question

Were fictional parsing, normalization, masking, aggregation, or enrichment steps successful and current?

Strong fictional evidence

Parser version, transformation status, dropped fields, mapping tests, enrichment freshness, and defects.

Important limitation

Successful processing does not prove the original source was complete.

Duplication and uniqueness

Review question

Could fictional records represent the same underlying activity more than once or collapse multiple activities into one?

Strong fictional evidence

Source identifiers, event identifiers, correlation keys, duplicate rate, transformation, and sampling.

Important limitation

Removing duplicates incorrectly may hide meaningful repeated behavior.

Coverage

Review question

Which fictional identities, devices, services, environments, zones, states, and time periods are represented?

Strong fictional evidence

Coverage map, inventory comparison, expected sources, sampling, exclusions, and blind spots.

Important limitation

High overall coverage may still omit one critical service or identity class.

Access and availability

Review question

Can authorized fictional analysts and detections retrieve the needed fields within the required time?

Strong fictional evidence

Access role, query or retrieval health concept, storage status, latency, permissions, and owner.

Important limitation

Available evidence may still be too sensitive or unnecessary for the purpose.

Retention and deletion

Review question

Is fictional evidence available for the required detection and review period without exceeding approved privacy needs?

Strong fictional evidence

Purpose, retention period, archive state, deletion, access, legal or policy basis concept, and owner review.

Important limitation

Longer retention does not automatically improve detection quality.

Instructional Section 5

Review Eight Field Types

Directly recorded field

Fictional example

Fictional application operation result or identity role assignment.

Defensive value

May provide strong source-specific evidence when the schema and source are healthy.

Caution

The source may still record only one layer and may omit purpose or downstream impact.

Derived field

Fictional example

Fictional risk level calculated from identity role, asset value, and destination category.

Defensive value

Can simplify analyst understanding and detection logic.

Caution

The calculation, source versions, assumptions, and missing-data behavior must be documented.

Normalized field

Fictional example

Fictional common identity or action category mapped from different source schemas.

Defensive value

Supports comparison and reuse across source types.

Caution

Normalization may hide source-specific meaning or incorrectly merge distinct values.

Enriched field

Fictional example

Fictional service owner, device class, change state, or maintenance window.

Defensive value

Adds mission and authorization context.

Caution

Stale enrichment can create false confidence or wrong exclusions.

Aggregated field

Fictional example

Fictional count of approved requests in a defined time window.

Defensive value

Supports threshold and trend concepts without exposing every raw event.

Caution

Aggregation may hide sequence, uniqueness, or rare high-impact details.

Masked or grouped field

Fictional example

Fictional requester group rather than exact personal identity.

Defensive value

Reduces privacy exposure while preserving some defender questions.

Caution

Grouping may be too broad for privileged, supplier, or accountability decisions.

Missing field

Fictional example

Fictional destination-owner or approval-expiration value is absent.

Defensive value

The absence itself may identify a source, schema, lifecycle, or coverage problem.

Caution

Missing does not always mean false, unauthorized, or harmful.

Conflicting field

Fictional example

Fictional identity source shows role removed while group source still shows membership.

Defensive value

Supports reconciliation and source-health review.

Caution

The conflict may result from delay, scope, different authority, or mapping error.

Instructional Section 6

Build a Privacy-Aware Source Catalog

Purpose

State which fictional defender questions and decisions justify collecting the source.

Caution

Do not keep fields merely because they may be useful someday.

Field minimization

Collect fictional identity, device, service, operation, result, timing, and health fields only when needed.

Caution

Exact content or personal attributes may be unnecessary.

Role-based access

Limit fictional raw, enriched, sensitive, and source-health evidence to approved roles.

Caution

Broad access can reveal user activity and internal architecture.

Retention

Keep fictional evidence long enough for approved detection, testing, review, and recovery needs.

Caution

Longer retention creates privacy, access, storage, and misuse risk.

Masking and grouping

Use fictional requester groups, asset classes, service categories, and object categories when exact detail is unnecessary.

Caution

Grouping must not remove accountability for privileged or supplier decisions.

Portfolio separation

Use invented fields and records rather than sanitized real screenshots, logs, or source names.

Caution

Redaction can miss hidden sensitive details.

Deletion and retirement

Remove fictional data, fields, exports, and access when the purpose ends or the source is retired.

Caution

A retired detection does not automatically remove retained evidence.

Change review

Revalidate fictional privacy when fields, sources, enrichments, retention, users, services, or defender questions change.

Caution

A previously approved field set may become unnecessary or incomplete.

Instructional Section 7

Define Degraded-Source Decision States

Source stateFictional conditionDetection behaviorAnalyst guidance
HealthyRequired fields, freshness, completeness, timing, schema, transformation, and coverage meet approved expectations.Evaluate normal logic and confidence.Use standard triage while preserving normal evidence limits.
ConditionalOne noncritical field or enrichment is stale, but the core defender question can still be evaluated.Evaluate with reduced context or adjusted confidence.Avoid decisions that depend on the stale field.
DegradedOne required source, field, timing relationship, or coverage area is delayed, incomplete, or unreliable.Mark alerts provisional, reduce confidence, use alternate evidence, or limit logic according to design.Do not close or escalate high-impact conclusions without compensating evidence.
BlindRequired evidence is unavailable for a defined scope or period.Stop or separate unsupported logic, record the blind period, and avoid false Healthy status.Use approved alternate sources and reassess later when evidence is restored.
ConflictingAuthoritative and corroborating fictional sources disagree beyond expected delay or scope differences.Create a reconciliation condition rather than choosing one value silently.Review provenance, authority, timing, schema, transformation, and owner context.
RecoveringThe source has returned, but backlog, historical gaps, duplicate replay, clock, or schema validation remains incomplete.Use limited confidence until reconciliation and backfill status are understood.Reassess alerts created during the blind or degraded period.

Fictional Evidence Architecture

Northbridge Detection Data-Source Model

This conceptual model is completely invented and intentionally non-operational. It teaches source relationships without real platform names, log formats, fields, schemas, identities, events, domains, addresses, suppliers, or internal architecture.

Producing systems

Identity, endpoint, network, DNS, email, app, cloud

External context

Supplier, support, user confirmation, change, recovery

Collection

Exports, collectors, queues, timestamps, missing periods

Transformation

Parsing, normalization, masking, enrichment, aggregation

Fictional Detection Evidence Core

Questions

Mission, identity, service, destination, impact

Fields

Meaning, type, direct, derived, normalized, enriched

Timing

Event, collection, processing, alert, confirmation

Health

Freshness, completeness, queue, clock, schema

Coverage

Users, devices, services, zones, states, periods

Privacy

Purpose, minimization, access, retention, deletion

Decision

Confidence, alternatives, scope, impact, next evidence

Lifecycle

Owners, changes, tests, reviews, blind periods, retirement

Detection use

Logic, missing-data behavior, alert, confidence

Analyst use

Triage, correlation, alternatives, scope, impact

Owner use

Health, schema, privacy, changes, support, recovery

Portfolio boundary

Fully fictional, privacy-safe, non-operational

Fake Dashboard

Fake Northbridge Detection Source-Health Dashboard

Fictional source coverage, freshness, schema, privacy, ownership, and lifecycle status for training only.

Sources meeting full health requirements

8 / 12

Four fictional sources have delayed enrichment, schema drift, duplicate events, or incomplete coverage.

Sources with current field dictionaries

9 / 12

Three fictional sources changed fields or value meanings without completed documentation review.

Open evidence blind periods

3

Application correlation, supplier results, and historical support timing require bounded confidence.

Fake SOC Alert

Connected Network Source Shows Unexpected Evidence Drop

Source: Fake Northbridge Source Assurance Console • Time: 3:07 PM

High Severity
The fictional network collector reports Green connectivity after a parser update, but event volume falls below the expected range, one destination-owner field is missing, and application correlation is delayed. No evidence confirms event loss or harmful network behavior.
Defensive recommendation: Mark the fictional source Degraded. Review schema, parser version, field mapping, expected volume, queue age, clock, alternate sources, application correlation, affected detections, blind period, and rollback before restoring normal confidence.

Fake Log Panel

Fake Source Provenance and Health Timeline

training-log-viewer.log
09:00 SOURCE network-stream='connected'
09:08 CHANGE parser-version='updated'
09:16 VOLUME network-events='below-expected'
09:24 FIELD destination-owner='missing'
09:32 SOURCE application-correlation='delayed'
09:40 QUEUE network='normal'
09:48 CLOCK network='aligned'
09:56 SCHEMA validation='conditional'
10:04 TRANSFORM mapping='under-review'
10:12 COVERAGE affected-detections='4'
10:20 PRIVACY field-set='unchanged'
10:28 ALTERNATE firewall-events='current'
10:36 ALTERNATE application-events='delayed'
10:44 STATUS source='degraded'
10:52 CONFIDENCE network-observation='moderate'
11:00 CONFIDENCE application-outcome='low'
11:08 BLIND-PERIOD start='09:16'
11:16 OWNER source='assigned'
11:24 CONFIDENCE source-health='moderate'
15:07 ALERT issue='evidence-volume-drop'

Training note: this is fake data for defensive analysis practice only.

Fictional Evidence Matrix

What Source Evidence Supports—and What It Does Not Prove

SRC-01

Fictional identity-event stream

Observation

Role-assignment events are current, but group-membership updates arrive eight minutes later on average.

Supports

Role and effective-group state have different freshness characteristics.

Does not prove

The delay does not prove access remained active or that events were lost.

Detection-source use

Define separate confidence and delayed-group behavior for identity detections.

SRC-02

Fictional network sensor

Observation

The sensor reports Green connectivity and current heartbeat, but event volume falls below the expected range after a parser update.

Supports

Source availability and evidence completeness may differ.

Does not prove

Lower volume does not prove missing events, attack activity, or normal quiet behavior.

Detection-source use

Review schema, parser, field mapping, expected volume, and alternate sources.

SRC-03

Fictional DNS evidence

Observation

Resolver events are current, while application correlation for one migrated service is delayed.

Supports

Naming observations may be timely while service-outcome context remains incomplete.

Does not prove

Current DNS evidence does not prove the destination or application is correct or healthy.

Detection-source use

Lower conclusion confidence until application evidence catches up.

SRC-04

Fictional application audit stream

Observation

Required operation and result fields are present, but object-owner enrichment is two days old.

Supports

Core application evidence is current while contextual ownership may be stale.

Does not prove

Stale enrichment does not prove the operation was unauthorized.

Detection-source use

Avoid owner-based suppression or escalation until enrichment is refreshed.

SRC-05

Fictional supplier result feed

Observation

Requests and results are available, but duplicate result records appear after a delivery retry.

Supports

Counting logic may overstate supplier activity unless duplication is understood.

Does not prove

Duplicate records do not prove duplicate business processing.

Detection-source use

Document event identifiers, retry behavior, correlation, and safe deduplication tests.

SRC-06

Fictional support-ticket source

Observation

User-impact categories are available, but exact confirmation time is missing from older records.

Supports

The source can help with impact but may not support precise sequencing for the full retention period.

Does not prove

Missing timing does not prove the support record is invalid.

Detection-source use

Limit historical time-sequence claims and record the coverage boundary.

SRC-07

Fictional cloud administrative stream

Observation

The provider changed one action field from a detailed value to a broader category.

Supports

Existing logic and documentation may no longer represent the same behavior.

Does not prove

The schema change does not prove detection failure until tests are performed.

Detection-source use

Trigger field-dictionary review, regression testing, version update, and owner approval.

SRC-08

Fictional source-health service

Observation

The source-health dashboard is current, but its own storage status has not been independently validated.

Supports

Health evidence also requires provenance and dependency review.

Does not prove

The missing validation does not prove the health dashboard is wrong.

Detection-source use

Add independent confirmation and disclose residual confidence limits.

Analyze the Evidence

Which Source-Health Decision Is Best Supported?

The collector is connected and its heartbeat is current.
Event volume fell below the expected range after a parser update.
One destination-owner field is missing.
Application correlation is delayed.
Queue age and clock alignment are normal.
Firewall evidence is current and can provide limited corroboration.
No supplied evidence confirms event loss or harmful network behavior.
Four detections depend on the affected field or correlation.

Which conclusion most responsibly represents the fictional network-source evidence?

Common Mistakes

Avoid Ten Data-Source Errors

Source exists, so coverage is complete

Fictional observation

A fictional identity feed is listed in the catalog and treated as complete for every user and service.

Decision impact

Excluded environments, service identities, delayed groups, or recovery roles may be missed.

Professional correction

Document scope, expected populations, exclusions, freshness, fields, and blind periods.

Connected means healthy

Fictional observation

A fictional network collector is Green while event freshness and volume are degraded.

Decision impact

Detections may run with stale or incomplete evidence.

Professional correction

Measure connectivity, freshness, completeness, queue, clock, schema, transformation, and coverage separately.

Field name equals field meaning

Fictional observation

A fictional field called result is assumed to represent business success across every source.

Decision impact

Logic may compare values that have different meanings.

Professional correction

Use a field dictionary with source, version, type, meaning, values, owner, and limitations.

Normalization removes differences safely

Fictional observation

Fictional identity and action values from several sources are merged into broad categories.

Decision impact

Source-specific distinctions may disappear or be mapped incorrectly.

Professional correction

Preserve original source context and test normalized mappings.

More data is always better

Fictional observation

A fictional detection collects detailed personal, message, or object data unrelated to the defender question.

Decision impact

Privacy, access, retention, and analyst-overload risk increases.

Professional correction

Use purpose-based field minimization and role-limited access.

One source proves the full outcome

Fictional observation

A fictional network event is treated as proof that an application operation succeeded.

Decision impact

Transport, application authorization, object state, and user outcome may be confused.

Professional correction

Correlate network, identity, application, support, and source-health evidence.

Duplicates are simply deleted

Fictional observation

A fictional supplier feed contains repeated events and all repeated records are discarded.

Decision impact

Legitimate repeated activity or retry behavior may be hidden.

Professional correction

Understand identifiers, retry semantics, aggregation, and test deduplication carefully.

Missing field equals malicious behavior

Fictional observation

A fictional approval field is absent and the event is labeled unauthorized.

Decision impact

Schema, collection, transformation, delay, or coverage problems may be misclassified.

Professional correction

Treat missing data as an evidence condition requiring validation.

Health source is assumed perfect

Fictional observation

A fictional source-health dashboard is trusted without reviewing its own dependencies and freshness.

Decision impact

The program may gain false confidence about evidence quality.

Professional correction

Validate source-health provenance, availability, storage, clock, and independent checks.

Real logs appear in a learning artifact

Fictional observation

A fictional portfolio includes copied internal fields, screenshots, domains, event values, supplier records, or user activity.

Decision impact

Sensitive systems, people, and defensive capabilities may be exposed.

Professional correction

Invent every source, field, event, value, identity, owner, date, and outcome.

Safe Fictional Practice Lab

Build the Northbridge Detection Data-Source Catalog

Use only the supplied fictional information on this page. Do not collect, query, export, inspect, monitor, search, correlate, copy, test, or modify any real telemetry, account, endpoint, network, domain, application, cloud service, supplier, platform, source, schema, field, or organization.
1

Define defender questions

List the fictional identity, endpoint, network, DNS, email, application, cloud, supplier, administrative, support, and source-health questions the program must answer.

Required output

Defender-question and evidence-needs catalog.

Quality check

Each question names one decision and one non-proof statement.

2

Create the source inventory

Document fictional source category, owner, producing system class, purpose, scope, environments, fields, retention, access, and dependencies.

Required output

Detection data-source catalog.

Quality check

No source is described only by a product or platform label.

3

Build the field dictionary

Define fictional field name, source, version, type, meaning, allowed values, direct or derived state, privacy, owner, and limits.

Required output

Detection field dictionary.

Quality check

Field meaning remains source-specific where necessary.

4

Map provenance

Trace fictional event generation, collection, parsing, normalization, enrichment, storage, detection, alert, analyst decision, and feedback.

Required output

Evidence-provenance map.

Quality check

Every transformation and owner is visible.

5

Define source-health measures

Record fictional connectivity, freshness, completeness, clock, schema, transformation, duplication, coverage, access, retention, and blind-period measures.

Required output

Source-health requirements matrix.

Quality check

Green connectivity alone cannot produce a Healthy rating.

6

Review privacy and access

Specify fictional purpose, required fields, minimization, analyst roles, retention, deletion, sharing, and portfolio exclusions.

Required output

Detection evidence privacy plan.

Quality check

Every field is justified by an approved defender question.

7

Identify gaps and conflicts

Document fictional missing sources, delayed fields, conflicting values, stale enrichment, duplication, incomplete populations, and unsupported time periods.

Required output

Coverage-gap and evidence-conflict register.

Quality check

Gaps are not converted into unsupported behavior conclusions.

8

Define degraded-source behavior

State how fictional detection confidence, severity, analyst guidance, alternate evidence, suppression, and closure change during source degradation.

Required output

Degraded-evidence decision matrix.

Quality check

The design avoids both silent failure and false certainty.

9

Create review triggers

Assign fictional source, schema, parser, field, application, identity, supplier, privacy, retention, and owner change triggers.

Required output

Source lifecycle and recertification plan.

Quality check

Each trigger has an owner, due date, validation, and documentation update.

10

Assemble the portfolio package

Combine the fictional catalog, field dictionary, provenance, health, privacy, gaps, degraded behavior, owners, residual risks, and executive summary.

Required output

Detection data-source governance package.

Quality check

The final artifact is traceable, maintainable, privacy-safe, and fully fictional.

Scenario Decision Lab

The Best Source Is Missing One Required Population

A fictional identity source is current and well documented for employees, but service identities and recovery roles are excluded. A proposed detection is intended to cover all privileged identities.

Scenario Decision Lab

A Schema Change Preserves Records but Changes Meaning

A fictional cloud source continues sending records after an update. One detailed administrative-action field is replaced by a broad category, but existing logic and documentation still assume the old meaning.

Advanced Challenge

Design Source Governance for a Multi-Source Detection Program

Fictional Northbridge wants detections across identity, endpoint, network, DNS, application, cloud, supplier, support, and recovery workflows. The sources have different owners, schemas, delays, populations, privacy needs, retention periods, transformations, and failure modes. Leadership assumes that collecting all of them will automatically provide strong detection coverage.

Build question-source mapping

Connect each fictional defender question to primary, corroborating, enrichment, and health sources.

Create provenance standards

Document fictional generation, collection, parsing, normalization, enrichment, storage, detection, alert, and feedback.

Define health states

Use fictional Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering decisions.

Measure coverage

Map fictional identities, devices, services, environments, states, time periods, exclusions, and blind spots.

Protect privacy

Justify fictional fields, access, retention, masking, deletion, sharing, and portfolio separation.

Control change

Trigger fictional review after schema, parser, platform, source, application, identity, supplier, privacy, or owner change.

Challenge output

Produce a fictional source-governance charter, defender-question mapping, source inventory, field dictionary, provenance map, health-state model, privacy plan, retention matrix, coverage-gap register, degraded-source guidance, owner matrix, change-control process, residual-risk statement, and leadership summary.

Defender Habits

Data Sources for Detection Checklist

Check Your Understanding

A5.2 Mini Quiz: Data Sources for Detection

Choose your answers first. Explanations appear only after submission.

1. What is the strongest way to select a fictional detection data source?

2. A fictional collector reports Green connectivity, but its last event is twenty minutes old. What is the strongest conclusion?

3. Why is a field dictionary important?

4. A fictional application source is current, but owner enrichment is stale. What is safest?

5. Why can duplicate fictional records be difficult?

6. Which privacy approach is strongest?

7. Which portfolio approach is safest?

Portfolio Prompt

Portfolio Prompt

Create a fully fictional Detection Data-Source Governance Package for the Northbridge Student-Support Cooperative. Include mission, purpose, stakeholders, scope, exclusions, safety boundary, at least thirty defender questions, primary sources, corroborating sources, enrichment sources, source-health sources, identity evidence, endpoint evidence, network evidence, DNS evidence, email evidence, application evidence, cloud evidence, supplier evidence, administrative evidence, support evidence, source owners, schemas, field dictionaries, direct fields, derived fields, normalized fields, enriched fields, aggregated fields, grouped fields, missing fields, conflicting fields, event time, collection time, processing time, alert time, confirmation time, recovery time, provenance chains, parsing, normalization, enrichment, storage, access, retention, deletion, privacy, coverage, populations, environments, states, time periods, exclusions, blind spots, blind periods, connectivity, freshness, completeness, clock, queue age, schema health, transformation health, duplication, access availability, Healthy states, Conditional states, Degraded states, Blind states, Conflicting states, Recovering states, alternate evidence, confidence changes, affected detections, change notifications, review triggers, residual risks, leadership summary, analyst guide, reflection, and a statement that every organization, source, field, schema, event, identity, owner, date, decision, and outcome is invented.

Start with fictional defender questions before deciding which sources or fields to collect.
Treat provenance, field meaning, timing, health, coverage, privacy, and ownership as part of detection quality.
Separate source connectivity from event freshness, completeness, semantic meaning, and decision usefulness.
Define degraded-source behavior so detections do not fail silently or create false certainty.
Keep the entire artifact completely fictional, defensive, non-operational, privacy-safe, evidence-aware, maintainable, and suitable for a public learning portfolio.

Confidence / Readiness Reflection

Are You Ready for Detection Logic Concepts?

Before moving to A5.3, rate your readiness from 1 to 5 for source categories, defender questions, provenance, field meaning, timing, source health, coverage, transformations, duplication, privacy, degraded states, ownership, review triggers, and complete fictionalization.

I can explain why one fictional source rarely proves a complete identity, service, or business outcome.
I can trace evidence from producing system through alert and analyst decision.
I can distinguish event, collection, processing, alert, confirmation, and recovery time.
I can separate connectivity, freshness, completeness, schema, transformation, and coverage.
I can explain how stale enrichment or conflicting sources change confidence.
I can design privacy-aware field selection and retention.
I can define detection behavior during Degraded or Blind source states.
I can produce a safe fictional source catalog without copying real telemetry or internal details.
Record one fictional defender question, one primary source, one corroborating source, one required field, one provenance risk, one source-health measure, one privacy decision, and one question you will carry into A5.3.

Key Takeaways

What You Should Remember

1.Detection data sources should be selected according to fictional defender questions, not volume, popularity, or convenience.
2.Identity, endpoint, network, DNS, email, application, cloud, supplier, administrative, support, and source-health evidence answer different questions and have different limitations.
3.Provenance connects fictional event generation, collection, parsing, normalization, enrichment, storage, detection, alerting, analyst decisions, and feedback.
4.Field names do not guarantee field meaning; source, schema, version, type, transformation, ownership, and limitations matter.
5.Event time, collection time, processing time, alert time, confirmation time, and recovery time support different conclusions.
6.Connectivity, freshness, completeness, clock, schema, transformation, duplication, coverage, access, and retention are separate source-health dimensions.
7.Missing, delayed, stale, duplicated, or conflicting fictional evidence should change confidence and guidance rather than become proof of harmful behavior.
8.Purpose-based field minimization, access, retention, deletion, masking, and portfolio separation protect privacy.
9.Source governance requires owners, health states, blind-period records, change notifications, review triggers, validation, and retirement.
10.Every CyberShield detection-source artifact must remain fully fictional, authorized, defensive, non-operational, privacy-safe, and incapable of exposing real systems or people.

Navigation

Continue Module A5

Next, translate fictional defender questions and evidence into conceptual detection logic using conditions, sequences, counts, time windows, relationships, context, exclusions, source-health states, severity, confidence, and missing-data behavior.