High School AdvancedModule A6Lesson 2 of 10Provenance, Schemas, Timing, Quality, and Source Health

A6.2 Log Collection and Normalization Concepts

Learn how fictional evidence moves from source systems through selection, collection, parsing, field mapping, normalization, enrichment, storage, indexing, and analyst use—while preserving source meaning, timing, health, privacy, ownership, and limitations.

Lesson Progress

Log Collection and Normalization Concepts

High School AdvancedA6: SIEM and Alert Triage Concepts • Lesson 2 of 10

20% complete

Readiness Check

Before You Start

0/6 ready

Professional Hook

The Same Normalized Value Can Hide Different Source Meanings

A fictional identity source reports a request as accepted, while a fictional application source reports a workflow as completed. Both values are normalized to Success. A correlation then treats them as equivalent, even though one source records request acceptance and the other records business completion. The fields look consistent, but the meaning is not.

Weak interpretation

“Both normalized fields say Success, so the same outcome occurred.”

Strong interpretation

“The shared category supports comparison, but source meaning, transformation, timing, and limitations remain necessary before the values are treated as equivalent.”

Normalization creates consistency. It should not erase provenance, uncertainty, or meaningful source differences.

Exactly Five Learning Objectives

What You Will Be Able to Do

Objective 1

Trace a fictional record from source creation through collection, parsing, normalization, enrichment, storage, indexing, correlation readiness, and analyst use while preserving provenance and timing.

Objective 2

Differentiate fictional raw evidence, parsed fields, normalized fields, transformed values, enrichment, derived context, source-health metadata, and analyst interpretation.

Objective 3

Evaluate fictional collection and normalization quality using schema versions, field meaning, event time, collection time, processing time, delay, duplication, loss, ordering, coverage, privacy, retention, and ownership.

Objective 4

Design safe fictional source-health behavior for Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering collection states.

Objective 5

Create a portfolio-ready fictional Collection and Normalization Package containing a source catalog, data dictionary, pipeline map, quality controls, failure modes, validation cases, metrics, owners, limitations, and review triggers.

Why This Matters

Collection and Normalization Shape Every Later SIEM Decision

Fictional searches, correlations, alerts, dashboards, cases, metrics, and leadership decisions depend on the records and fields available in the SIEM. A missing population, delayed source, parser defect, stale schema, incorrect mapping, duplicate path, privacy-heavy field set, or hidden transformation can affect confidence and coverage across many detections at once.

Evidence integrity

Preserve fictional source meaning, timing, provenance, uniqueness, completeness, and transformation history.

Decision integrity

Prevent fictional missing, delayed, duplicated, or semantically changed evidence from becoming false confidence.

Lifecycle integrity

Maintain fictional schemas, parsers, mappings, owners, tests, metrics, privacy, retention, review triggers, and retirement.

Core Framework

The P-I-P-E-L-I-N-E Method

P — Purpose-limit collection

Select fictional sources and fields because they support documented defender questions, privacy needs, access roles, retention, and lifecycle.

I — Identify provenance

Preserve fictional source category, record identity, schema, parser, event time, collection path, owner, and local meaning.

P — Parse with version control

Interpret fictional records using documented schemas, parser versions, Unknown handling, failure behavior, and tests.

E — Establish field mappings

Connect fictional source fields to shared fields while retaining source values, transformations, assumptions, and limitations.

L — Label normalization layers

Distinguish fictional raw evidence, parsed fields, normalized fields, enrichment, derived context, and analyst interpretation.

I — Inspect timing and uniqueness

Review fictional event, collection, processing, indexing, and alert times plus retries, replays, duplicates, and order.

N — Note source health

Use fictional Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering states.

E — Evaluate quality and coverage

Measure fictional completeness, semantic quality, freshness, failures, duplicates, privacy, coverage, debt, and residual risk.

Decision-ready pipeline statement

This fictional collection and normalization design preserves source purpose, provenance, schema, parser, source field, normalized field, transformation, timing, source health, uniqueness, coverage, privacy, ownership, quality, limitations, review triggers, and lifecycle.

Advanced Vocabulary

Terms for Collection and Normalization

Source record

A fictional original record produced by an identity, endpoint, network, DNS, email, application, cloud, supplier, change, support, administrative, or source-health system.

Raw evidence

A fictional representation of source-created information before SIEM parsing, normalization, enrichment, or analyst interpretation.

Collector

A fictional defensive component or process that receives, transfers, or forwards selected records from a source into the evidence pipeline.

Collection scope

The fictional identities, services, devices, environments, fields, states, and periods that a collector is designed to include.

Parsing

The fictional interpretation of a source record into documented source-specific fields and values.

Parser version

The fictional version of parsing logic that defines how a source record is interpreted.

Schema

A fictional documented structure describing fields, types, allowed values, relationships, and meanings.

Schema drift

A fictional change in source fields, types, values, relationships, or meaning that may affect parsing and normalization.

Field mapping

The fictional relationship between a source-specific field and a shared normalized field.

Normalization

The fictional mapping of different source fields into common categories while preserving provenance, source meaning, and limitations.

Canonical field

A fictional shared field used across multiple source categories for consistent defensive analysis.

Transformation

A fictional documented change to a value, type, timestamp, category, format, or relationship during processing.

Enrichment

Fictional identity, device, service, ownership, authorization, criticality, destination, peer, change, policy, or mission context added after collection.

Derived field

A fictional field calculated or inferred from one or more source fields rather than directly recorded by the source.

Event time

The fictional time when the underlying activity or state occurred at the source.

Collection time

The fictional time when the collector received or transferred the record.

Processing time

The fictional time when parsing, normalization, enrichment, storage, indexing, or correlation preparation occurred.

Ingestion delay

The fictional difference between event time and the time evidence becomes available in the SIEM.

Duplicate record

A fictional repeated representation of the same underlying event, request, retry, replay, state, or collector delivery.

Out-of-order arrival

A fictional condition in which records enter the SIEM in a different order from the underlying event sequence.

Data loss

A fictional gap in which expected records, fields, populations, environments, or periods are unavailable.

Field completeness

The fictional degree to which required and optional fields are present for a defined source and purpose.

Semantic quality

The fictional degree to which field values and categories retain their intended meaning after processing.

Normalization debt

Fictional risk created by stale mappings, missing tests, undocumented transformations, weak ownership, or unresolved schema differences.

Instructional Section 1

Trace Ten Collection and Normalization Stages

1. Source creation

A fictional source records an activity, state, transaction, result, ownership change, support action, or source-health observation.

Fictional evidence

Original source category, schema version, record identifier, event time, field values, local meaning, and source owner.

Risk

The original record may already be incomplete, delayed, duplicated, ambiguous, or outside intended coverage.

Professional control

Document source purpose, local field meanings, expected record behavior, owner, schema, and limitations.

2. Purpose-based selection

The fictional program decides which records, fields, identities, services, environments, and periods should enter the SIEM.

Fictional evidence

Collection scope, field-purpose map, privacy review, access roles, retention need, exclusions, and owner approval.

Risk

Overcollection increases privacy and maintenance risk while undercollection creates coverage gaps.

Professional control

Use mission-driven selection, field minimization, documented exclusions, and periodic coverage review.

3. Collection

A fictional collector receives or transfers selected records from the source.

Fictional evidence

Collection time, collector identifier, delivery path, queue state, retry state, expected volume, and source health.

Risk

Records may be delayed, lost, duplicated, retried, reordered, or delivered through overlapping paths.

Professional control

Track freshness, completeness, queue age, delivery success, uniqueness, blind periods, and recovery.

4. Parsing

A fictional parser interprets source content into source-specific fields.

Fictional evidence

Parser version, schema version, extraction status, parse errors, unmapped values, and fallback behavior.

Risk

A parser may silently drop fields, misread values, or continue after a source schema change.

Professional control

Use versioned schemas, parser tests, parse-failure alerts, Unknown handling, and owner review.

5. Field mapping

Fictional source-specific fields are mapped to shared normalized categories.

Fictional evidence

Source field, canonical field, type, allowed values, transformation, confidence, provenance, and owner.

Risk

Different source meanings may be treated as equivalent.

Professional control

Preserve source-specific meaning, document assumptions, and test semantic differences.

6. Normalization

Fictional records are organized into consistent defensive structures for search, dashboards, correlation, and triage.

Fictional evidence

Canonical fields, normalized values, source labels, timestamps, missing-data states, and transformation history.

Risk

Normalization can hide source nuance, time differences, missing fields, or uncertainty.

Professional control

Retain provenance, source category, local values, transformation notes, and limitations.

7. Enrichment

Fictional identity, device, service, ownership, authorization, peer, destination, criticality, change, policy, or mission context is added.

Fictional evidence

Enrichment source, owner, timestamp, freshness, relationship, privacy purpose, and limitation.

Risk

Stale or incorrect enrichment may change severity, priority, ownership, or interpretation.

Professional control

Track freshness, authority, owner, expiration, confidence, and alternate context.

8. Storage and indexing

Fictional processed records become available for approved search, correlation, dashboard, case, and quality purposes.

Fictional evidence

Storage state, index status, retention, access roles, availability, deletion, and search scope.

Risk

Missing results may reflect indexing, retention, access, or availability rather than absence.

Professional control

Document searchable fields, retention boundaries, access limits, index health, and deletion behavior.

9. Correlation readiness

Fictional records are prepared for relationship, count, sequence, threshold, state, timing, and source-health analysis.

Fictional evidence

Correlation keys, normalized identities, session and service relationships, destination categories, time alignment, and uniqueness.

Risk

Incorrect relationships or timing may create false matches or missed conditions.

Professional control

Validate keys, windows, uniqueness, time basis, missing-data behavior, source health, and test coverage.

10. Analyst use and lifecycle

Fictional analysts use processed evidence for searches, alerts, dashboards, cases, decisions, quality review, change, and retirement.

Fictional evidence

Alert explanations, case notes, owner questions, health visibility, quality findings, change records, and review triggers.

Risk

Processed evidence may be treated as unquestioned fact or remain stale after source changes.

Professional control

Train analysts on provenance and limits, assign owners, measure quality, review changes, and retire outdated mappings.

Instructional Section 2

Normalize Eight Fictional Source Categories

Identity and access

Fictional records

Role state, group state, approval, extension, sponsor, session, assignment, revocation, and source-health records.

Normalization needs

Identity category, role category, authorization state, approval timing, session relationship, owner group, and lifecycle state.

Semantic risk

A valid identity may be mistaken for valid use, or role assignment may be mistaken for effective access.

Privacy boundary

Collect purpose-limited identity, role, owner, and authorization fields rather than unrelated personal history.

Endpoint and device

Fictional records

Device identity, device class, onboarding, owner, replacement, posture, support state, network class, and retirement.

Normalization needs

Device category, managed state, owner group, lifecycle state, support status, and source health.

Semantic risk

A device label may not prove who used the device or whether inventory is current.

Privacy boundary

Use device and ownership categories without personal content.

Network

Fictional records

Source group, destination class, direction, session, policy result, service relationship, timing, and sensor health.

Normalization needs

Source zone, destination zone, direction, relationship type, session state, policy result, and time basis.

Semantic risk

A destination relationship does not prove application action, content, owner, or harmful intent.

Privacy boundary

Use service and destination categories rather than real addresses or browsing history.

DNS and naming

Fictional records

Requester group, resolver, question category, response category, cache state, policy, timing, and health.

Normalization needs

Requester class, resolver class, question type, response class, audience, cache state, and policy result.

Semantic risk

Different answers may reflect audience, cache, migration, policy, or timing rather than harmful behavior.

Privacy boundary

Limit evidence to defined service questions rather than broad personal query histories.

Application and service

Fictional records

Operation category, result, object class, service owner, request, response, state, user impact, and source health.

Normalization needs

Operation, result, object class, service category, owner group, user-impact state, and workflow state.

Semantic risk

The same word such as success may represent different outcomes across applications.

Privacy boundary

Use object and result categories rather than message contents or personal records.

Cloud and administration

Fictional records

Administrative role, configuration change, service action, policy result, owner, approval, timing, and source health.

Normalization needs

Administrative action category, privilege class, change state, approval state, owner group, and service scope.

Semantic risk

An administrative action may be expected, automated, emergency, or incomplete depending on context.

Privacy boundary

Exclude real configuration detail, credentials, routes, and unnecessary operator information.

Supplier and support

Fictional records

Supplier identity, sponsor, assignment, device, destination, session, support request, result, closure, and source health.

Normalization needs

Supplier role, sponsor state, assignment state, device class, support purpose, destination class, and closure state.

Semantic risk

Outside-hours activity may be approved maintenance or emergency support rather than unauthorized behavior.

Privacy boundary

Use role, sponsor, purpose, destination category, and assignment timing rather than personal supplier details.

Change and recovery

Fictional records

Change identifier, owner, scope, expected behavior, start, end, validation, result, rollback, recovery, and closure.

Normalization needs

Change state, operating state, approved scope, expected difference, rollback state, and completion state.

Semantic risk

A linked change does not authorize every identity, destination, action, result, or time.

Privacy boundary

Avoid configuration and architecture details unnecessary for the decision.

Instructional Section 3

Build a Fictional Field Dictionary

Normalized fieldSource examplesMeaningTransformationLimitation
actor.categoryFictional user, service, supplier, administrative, or recovery identity labels.Shared category describing the kind of actor represented by the record.Mapped into invented User, Service, Supplier, Administrative, Recovery, or Unknown categories.Actor category does not prove authorization, ownership, or intent.
action.categoryFictional sign-in, role change, service operation, destination request, support action, or policy evaluation.Shared category describing observed activity or state transition.Mapped from source-specific action names into conceptual defensive categories.Different source actions may share a category while retaining different meanings.
result.categoryFictional completed, accepted, allowed, denied, pending, failed, partial, or unknown values.Shared category describing the source-reported outcome.Normalized into invented Success, Denied, Failed, Pending, Partial, or Unknown values.Success in one source may mean request acceptance rather than business completion.
service.categoryFictional student-support, notification, identity, recovery, administrative, learning, or supplier-service labels.Shared mission category for the associated service.Mapped from source-specific names into public-safe fictional classes.Service category may be stale after ownership, architecture, or mission changes.
destination.categoryFictional approved dependency, new dependency, supplier service, recovery service, or unknown destination.Shared category describing the destination relationship.Derived from source evidence and an invented service-relationship catalog.Destination category does not prove application result or content.
authorization.stateFictional approval, assignment, extension, sponsor, maintenance, or emergency-use records.Shared state describing whether current authorization evidence supports the activity.Derived into Approved, Expired, Missing, Conflicting, or Unknown states.Authorization may depend on delayed or incomplete evidence.
event.timeFictional source-recorded activity or state-change time.Time when the underlying source event occurred.Converted to one fictional common format while preserving source time and clock state.Event time may be missing, rounded, delayed, or affected by clock differences.
source.healthFictional freshness, completeness, schema, parser, queue, clock, coverage, duplication, blind-period, and recovery checks.Shared state describing how reliably evidence supports its documented purpose.Derived into Healthy, Conditional, Degraded, Blind, Conflicting, or Recovering.Healthy technical status does not guarantee perfect semantic quality.

Instructional Section 4

Use Six Timing Models

Event-time reasoning

Use

Understand the fictional order and duration of underlying activities or state changes.

Defender question

When did the source say the activity happened?

Risk

Source clocks may be inaccurate, rounded, missing, or differently synchronized.

Collection-time reasoning

Use

Understand when fictional evidence entered the collection pipeline.

Defender question

When did the collector receive or transfer the record?

Risk

Collection order may differ from event order.

Processing-time reasoning

Use

Understand when fictional parsing, normalization, enrichment, indexing, and correlation became available.

Defender question

When could the SIEM and analyst actually use the evidence?

Risk

Processing delay may create late alerts or incomplete cross-source context.

Alert-time reasoning

Use

Understand when the fictional correlation or alert condition was evaluated and presented.

Defender question

When did the platform produce the alert relative to the event?

Risk

Alert time may be mistaken for event time.

Window reasoning

Use

Determine whether fictional records belong in a count, sequence, expiration, or correlation window.

Defender question

Which time basis and tolerance define the window?

Risk

A wrong time basis may create false matches or missed conditions.

Recovery reasoning

Use

Understand fictional blind periods, backlog, replay, duplication, historical reassessment, and restoration confidence.

Defender question

Which records are original, delayed, replayed, duplicate, missing, or unreconciled?

Risk

Source connectivity may be mistaken for complete evidence recovery.

Instructional Section 5

Use Six Source-Health States

Healthy

Condition

Required fictional records and fields are current, complete enough, correctly parsed, correctly mapped, aligned, covered, and accessible.

Pipeline behavior

Normal collection, parsing, normalization, enrichment, storage, indexing, and analyst use may proceed.

Analyst meaning

Use normal confidence while preserving ordinary semantic and coverage limits.

Conditional

Condition

One optional field, enrichment source, owner record, or noncritical relationship is stale or incomplete.

Pipeline behavior

Preserve the core record but mark enrichment-dependent severity, priority, or routing as limited.

Analyst meaning

Do not use stale context for closure or broad conclusions.

Degraded

Condition

A required source, field, parser, schema, mapping, clock, queue, or collection path is delayed or incomplete.

Pipeline behavior

Make affected records and fields visibly limited, lower confidence, and request alternate evidence.

Analyst meaning

Do not interpret missing or delayed evidence as normal-confidence absence.

Blind

Condition

Required evidence is unavailable for a documented population, service, environment, or period.

Pipeline behavior

Record the blind period, affected sources and detections, alternate evidence, and reassessment need.

Analyst meaning

Do not claim the condition was absent.

Conflicting

Condition

Source, parser, schema, mapping, timing, or ownership evidence disagrees beyond expected differences.

Pipeline behavior

Preserve both versions, show provenance, and create a reconciliation state.

Analyst meaning

Avoid silently trusting one value without authority and timing review.

Recovering

Condition

The pipeline is available again, but backlog, replay, duplication, schema, clock, field, or historical gaps remain.

Pipeline behavior

Use limited confidence until uniqueness, backfill, semantic quality, and regression checks pass.

Analyst meaning

Connectivity restoration does not prove the evidence pipeline is fully healthy.

Instructional Section 6

Validate Eight Collection Cases

Fictional caseInputExpected resultFailure may indicate
Expected recordA fictional identity role-state record arrives with required fields, current schema, and Healthy source status.The record parses, maps, normalizes, stores, and remains searchable with complete provenance.Possible collection, parser, mapping, indexing, or test-data defect.
Missing required fieldA fictional approval record arrives without approval_end.The record remains visible with missing-field state; authorization becomes Conditional or Unknown.Silent field loss or unsupported authorization confidence.
Missing optional fieldA fictional service record lacks criticality enrichment but includes action, result, owner, and source health.Core evidence remains usable while severity or priority context is limited.The pipeline depends unnecessarily on optional enrichment.
Schema driftA fictional source changes result_code to a new undocumented string.Unknown is preserved, a quality alert appears, and affected rules remain Conditional.Silent semantic corruption or false normalization.
Duplicate deliveryThe same fictional record identifier arrives through retry and replay paths.The duplicate relationship is visible and uniqueness rules avoid inflated counts without deleting legitimate repetition.Duplicate flooding or overaggressive deduplication.
Out-of-order arrivalA fictional revocation occurs first but reaches the SIEM after session evidence.Event-time and collection-time order remain distinct; sequence confidence reflects delay.Collection order is mistaken for event order.
Blind periodA fictional source provides no records for a documented period and population.Blind state and affected coverage are visible; quiet results are not labeled normal or absent.False-negative and false-confidence risk.
Recovery replayA fictional source returns and delivers queued records with original event times plus replay metadata.Recovering state, duplicate-aware handling, historical reassessment, and backfill validation.Duplicate alerts, false sequence, missed blind-period cases, or premature Healthy status.

Instructional Section 7

Measure Eight Collection Quality Dimensions

Collection completeness

Review question

Did the fictional pipeline receive expected records across documented identities, services, environments, and periods?

Fictional evidence

Expected volume, source counters, sequence checks, blind-period records, owner confirmation, and coverage maps.

Limitation

Expected volume may itself be inaccurate.

Field completeness

Review question

Are fictional required and optional fields present at expected rates?

Fictional evidence

Field-presence reports, schema tests, parser results, unknown-value counts, and source-owner review.

Limitation

A present field may still contain incorrect meaning.

Semantic quality

Review question

Do fictional normalized values preserve source meaning and documented differences?

Fictional evidence

Source-to-canonical mappings, test cases, owner review, unknown values, and correlation outcomes.

Limitation

Semantic quality can change after source workflow or policy changes.

Freshness and delay

Review question

How quickly do fictional records become available after event time?

Fictional evidence

Event, collection, processing, indexing, and alert timestamps with source states.

Limitation

Faster delivery does not guarantee correctness or completeness.

Duplicate rate

Review question

How often do fictional retries, replays, overlapping collectors, or continuing-state records create repeated representations?

Fictional evidence

Record identifiers, correlation keys, retry state, replay markers, collector paths, and grouping outcomes.

Limitation

Repeated records may represent legitimate repeated actions.

Parse and mapping failures

Review question

How often do fictional records fail parsing, produce Unknown values, or miss mappings?

Fictional evidence

Parser errors, unmapped fields, Unknown values, schema versions, fallback behavior, and affected detections.

Limitation

Low failure counts may hide silent misinterpretation.

Coverage confidence

Review question

Which fictional identities, services, devices, destinations, states, and periods have reliable collection and normalization?

Fictional evidence

Coverage map, source inventory, field dictionary, health states, tests, and known-gap register.

Limitation

Documented coverage does not prove every relevant behavior is observable.

Normalization debt

Review question

Which fictional mappings, parsers, schemas, tests, owners, transformations, and retirement tasks are stale or unresolved?

Fictional evidence

Debt register, review dates, owner matrix, change history, failed tests, and residual-risk records.

Limitation

Counting debt does not identify mission impact by itself.

Fictional Collection Architecture

Northbridge Source-to-SIEM Pipeline

This conceptual architecture is completely invented and intentionally non-operational. It teaches collection and normalization without real products, source names, schemas, fields, credentials, addresses, queries, records, systems, suppliers, or internal architecture.

Identity

Roles, approvals, groups, sessions, revocation

Device

Class, ownership, onboarding, replacement, support

Network and DNS

Relationships, direction, policy, naming, timing

Application and supplier

Actions, results, purpose, assignment, closure

Fictional Processing Pipeline

Select

Purpose, scope, privacy, fields, owners

Collect

Delivery, queues, retries, coverage, health

Parse

Schemas, fields, errors, versions, Unknowns

Map

Source fields, canonical fields, assumptions

Normalize

Shared categories with preserved provenance

Enrich

Identity, service, owner, change, mission context

Store

Indexing, access, retention, deletion, availability

Validate

Tests, metrics, defects, review, retirement

Search-ready evidence

Fields, provenance, timing, health, limits

Correlation-ready evidence

Keys, relationships, sequence, windows

Analyst-ready evidence

Observation, context, confidence, questions

Portfolio boundary

Fully fictional, privacy-safe, non-operational

Fake Dashboard

Fake Northbridge Collection Quality Dashboard

Fictional completeness, semantic quality, delay, duplication, parser status, coverage, privacy, and normalization debt for training only.

Source categories meeting collection gates

6 / 8

Supplier assignment and recovery identity coverage remain incomplete.

Canonical fields with validated semantic mappings

19 / 24

Five mappings require source-owner review after schema or workflow changes.

Open fictional pipeline defects

9

Delay, duplicates, Unknown values, stale enrichment, missing fields, privacy excess, indexing, recovery, and ownership remain open.

Fake SOC Alert

Normalized Evidence Quality Is Conditional

Source: Fake Northbridge Collection Quality Console • Time: 3:08 PM

High Severity
The fictional pipeline is receiving records, but one new result value maps to Unknown, two source meanings share the same normalized category, supplier assignment coverage is incomplete, recovery replay duplicates remain elevated, and unnecessary endpoint profile fields are still collected.
Defensive recommendation: Keep the fictional collection design Conditional. Resolve schema meaning, mapping semantics, source coverage, replay uniqueness, privacy minimization, owner assignment, validation tests, and review triggers before approval.

Fake Log Panel

Fake Collection and Normalization Timeline

training-log-viewer.log
09:02 SOURCE identity-event='created'
09:10 COLLECT identity-event='received'
09:11 PARSER version='identity-v3'
09:12 FIELD result-code='new-value'
09:12 PARSE result='unknown'
09:13 NORMALIZE actor-category='recovery'
09:13 NORMALIZE result-category='unknown'
09:14 ENRICH owner-group='identity-operations'
09:15 INDEX status='available'
09:16 HEALTH identity='degraded'
09:17 SOURCE supplier-assignment='coverage-gap'
09:18 SOURCE application='healthy'
09:19 DUPLICATE replay-rate='elevated'
09:20 PRIVACY endpoint-fields='excess'
09:21 MAPPING semantic-review='required'
09:22 TEST schema-drift='failed'
09:23 TEST recovery-replay='conditional'
09:24 DEBT normalization='9-open'
09:25 READINESS pipeline='conditional'
15:08 ALERT issue='normalized-evidence-quality'

Training note: this is fake data for defensive analysis practice only.

Fictional Evidence Matrix

What Collection Evidence Supports—and What It Does Not Prove

COLL-01

Fictional source catalog

Observation

Eight source categories are documented, but one supplier assignment feed and one recovery identity population remain out of scope.

Supports

Collection coverage is incomplete for some supplier and recovery questions.

Does not prove

The gap does not prove a missed event occurred.

Pipeline use

Keep affected detections Conditional and document residual risk.

COLL-02

Fictional processing timeline

Observation

An identity event occurs at 09:02, is collected at 09:10, processed at 09:13, and indexed at 09:15.

Supports

The evidence became searchable thirteen minutes after the event.

Does not prove

Delay does not prove the event is incorrect or harmful.

Pipeline use

Separate event, collection, processing, and availability time in analyst review.

COLL-03

Fictional parser report

Observation

A new source value appears in result_code and is mapped to Unknown instead of Success or Failure.

Supports

The parser preserved uncertainty rather than forcing a false category.

Does not prove

Unknown does not identify the correct business meaning.

Pipeline use

Escalate field meaning to the source owner and test affected correlations.

COLL-04

Fictional normalization review

Observation

Two application sources map completed and accepted to the same normalized Success value.

Supports

The canonical field may hide different workflow outcomes.

Does not prove

The mapping does not prove current alerts are wrong.

Pipeline use

Document source-specific semantics and refine mappings or alert explanations.

COLL-05

Fictional duplicate analysis

Observation

Three records share one event identifier but arrive through retry and recovery replay paths.

Supports

The records may represent one underlying event delivered multiple times.

Does not prove

Matching identifiers alone do not prove every repeated action is duplicate.

Pipeline use

Apply documented uniqueness and replay logic with regression tests.

COLL-06

Fictional source-health dashboard

Observation

Identity is Degraded, application is Healthy, DNS is Conditional, and network is Recovering.

Supports

Different evidence domains require different confidence and alternate-evidence decisions.

Does not prove

Health states do not prove the activities were expected or harmful.

Pipeline use

Expose source health in normalized records, alerts, dashboards, and cases.

COLL-07

Fictional privacy review

Observation

The proposed endpoint record includes personal profile fields unrelated to the device-class defender question.

Supports

The collection plan exceeds the documented purpose.

Does not prove

The finding does not invalidate all endpoint evidence.

Pipeline use

Remove unnecessary fields and retest analyst usefulness.

COLL-08

Fictional recovery report

Observation

The source is connected again, but replay backlog, duplicate rate, one schema change, and one missing period remain unresolved.

Supports

The pipeline is Recovering rather than Healthy.

Does not prove

The report does not prove every blind-period record is missing.

Pipeline use

Maintain limited confidence and historical reassessment until validation passes.

Analyze the Evidence

Which Collection Readiness Decision Is Best Supported?

Six of eight source categories meet current collection gates.
Supplier assignment and one recovery identity population remain outside reliable coverage.
One new result value is preserved as Unknown.
Two source meanings share one normalized Success category.
Recovery replay duplicates remain elevated.
Unnecessary endpoint profile fields are still collected.
Five canonical mappings require source-owner review.
Nine fictional pipeline defects remain open.

Which conclusion most responsibly represents the fictional Northbridge collection and normalization review?

Common Mistakes

Avoid Ten Collection and Normalization Errors

Raw, parsed, normalized, and enriched data are treated as identical

Fictional observation

A fictional analyst assumes a normalized field is the exact source record.

Decision impact

Transformations, mappings, missing fields, and derived context are hidden.

Professional correction

Label each evidence layer and preserve provenance, source field, transformation, and limitation.

Collection time is treated as event time

Fictional observation

A fictional sequence is built from the order records reached the SIEM.

Decision impact

Delayed or out-of-order evidence may create a false chronology.

Professional correction

Preserve event, collection, processing, indexing, and alert times separately.

Unknown values are forced into familiar categories

Fictional observation

A fictional new result value is automatically mapped to Success.

Decision impact

Schema drift becomes silent semantic corruption.

Professional correction

Preserve Unknown, alert on unmapped values, and obtain source-owner meaning.

All duplicate-looking records are deleted

Fictional observation

A fictional pipeline removes every repeated record identifier or field combination.

Decision impact

Legitimate repeated actions or state changes may disappear.

Professional correction

Define uniqueness, retry, replay, continuing-state, and break conditions with tests.

Normalization removes source context

Fictional observation

A fictional shared result field hides whether the source meant accepted, completed, or allowed.

Decision impact

Search and correlation may create false equivalence.

Professional correction

Preserve source category, local value, normalized value, mapping rationale, and limitation.

Healthy connectivity is treated as healthy evidence

Fictional observation

A fictional source is marked Healthy because records are arriving.

Decision impact

Schema, field, clock, completeness, semantic, or coverage defects may remain hidden.

Professional correction

Evaluate technical and semantic health together.

Every available field is collected

Fictional observation

A fictional source includes broad personal or operational detail unrelated to the defender question.

Decision impact

Privacy, access, retention, maintenance, and portfolio risk increase.

Professional correction

Use field-purpose mapping, minimization, role-based access, retention, deletion, and review.

Source recovery ends the issue

Fictional observation

A fictional source is marked Healthy immediately after reconnection.

Decision impact

Backlog, replay, duplicates, schema changes, missing periods, and false sequence remain unresolved.

Professional correction

Use Recovering until reconciliation and regression checks pass.

Mappings have no owner

Fictional observation

A fictional canonical field is used by many detections, but no owner reviews changes.

Decision impact

Semantic drift can affect multiple alerts silently.

Professional correction

Assign source, field, parser, normalization, detection, privacy, and lifecycle owners.

Real schemas or records enter the portfolio

Fictional observation

A fictional artifact includes copied real field names, layouts, timestamps, addresses, screenshots, or identifiers.

Decision impact

Sensitive systems, people, suppliers, and architecture may be exposed.

Professional correction

Invent every source, schema, field, record, timestamp, owner, alert, and outcome.

Safe Fictional Practice Lab

Build the Northbridge Collection and Normalization Package

Use only the supplied fictional information on this page. Do not collect, copy, sanitize, upload, inspect, query, test, replay, parse, normalize, correlate, search, or modify any real log, source, schema, field, account, endpoint, network, domain, service, supplier, SIEM, platform, or organization.
1

Define the collection purpose

Write fictional defender questions for identity, endpoint, network, DNS, application, supplier, change, recovery, and source-health evidence.

Required output

Collection-purpose matrix.

Quality check

Every source and field supports a documented defensive decision.

2

Create the source catalog

Document fictional source category, owner, schema, event types, fields, timing, coverage, privacy, retention, health, and lifecycle.

Required output

Versioned source inventory.

Quality check

Known gaps and out-of-scope populations remain visible.

3

Map the pipeline

Trace fictional source creation, selection, collection, parsing, mapping, normalization, enrichment, storage, indexing, and analyst use.

Required output

Collection and normalization architecture.

Quality check

Every stage includes owner, timing, transformation, failure behavior, and limitation.

4

Build the field dictionary

Define fictional source fields, canonical fields, types, values, transformations, provenance, requirement, privacy purpose, and limitations.

Required output

Source-to-canonical field dictionary.

Quality check

Source-specific semantics remain visible after normalization.

5

Define timing

Document fictional event, collection, processing, indexing, and alert time plus clock state, delay, and windows.

Required output

Timing and freshness model.

Quality check

Collection order is never treated automatically as event order.

6

Define source-health states

Create fictional Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering behavior for sources, fields, parsers, mappings, queues, and coverage.

Required output

Collection source-health model.

Quality check

Missing or degraded evidence cannot silently become normal confidence.

7

Create validation cases

Write fictional expected, missing-field, schema-drift, duplicate, out-of-order, blind-period, and recovery-replay tests.

Required output

Collection and normalization test library.

Quality check

Expected outcomes are documented before comparison.

8

Define quality metrics

Measure fictional completeness, field presence, semantic quality, freshness, delay, duplicates, parse failures, coverage, privacy, and debt.

Required output

Collection-quality metric dictionary.

Quality check

Each metric has definition, source, denominator, limitation, owner, action, and trigger.

9

Assign corrective actions

Connect fictional source, parser, mapping, timing, coverage, privacy, ownership, documentation, or retirement defects to owners, tests, rollback, and completion.

Required output

Collection defect and action register.

Quality check

Corrections address root causes rather than hiding evidence problems.

10

Prepare the portfolio package

Combine the fictional source catalog, pipeline, field dictionary, timing, health, tests, metrics, defects, owners, limitations, triggers, and reflection.

Required output

Public-safe Collection and Normalization Package.

Quality check

Every organization, source, schema, field, record, timestamp, owner, decision, and outcome is invented.

Scenario Decision Lab

A New Source Value Does Not Match the Existing Schema

A fictional application source begins sending accepted_pending in a field that previously contained only completed, denied, and failed. The current parser maps any unrecognized value to Success.

Scenario Decision Lab

A Source Reconnects after a Blind Period

A fictional source begins sending records again after a two-hour blind period. A backlog is replaying, duplicate delivery is elevated, one schema field changed, and a fifteen-minute period remains missing.

Advanced Challenge

Defend a Normalization Design before a Review Board

Fictional Northbridge wants one shared schema for identity, endpoint, network, DNS, application, cloud administration, suppliers, changes, and recovery. The current proposal has broad canonical fields but weak source provenance, unclear timing, unowned mappings, no Unknown-value policy, incomplete privacy review, no recovery tests, and no normalization-debt register.

Defend source selection

Explain which fictional defender questions justify each source, record type, field, access role, retention period, and privacy purpose.

Defend field meaning

Explain fictional source fields, canonical fields, transformations, semantic differences, Unknown values, and limitations.

Defend timing

Explain fictional event, collection, processing, indexing, and alert times plus delays, clocks, windows, and order.

Defend source health

Explain fictional Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering behavior.

Defend quality

Explain fictional completeness, semantic quality, duplicate rate, parse failures, coverage, privacy, and debt metrics.

Defend lifecycle

Explain fictional owners, tests, changes, review triggers, rollback, residual risk, mapping retirement, and source retirement.

Challenge output

Produce a fictional source catalog, source-to-canonical matrix, field dictionary, transformation register, timing model, source-health model, test library, quality dashboard, privacy review, ownership matrix, normalization-debt register, residual-risk statement, leadership summary, and public portfolio boundary.

Defender Habits

Log Collection and Normalization Checklist

Check Your Understanding

A6.2 Mini Quiz: Log Collection and Normalization Concepts

Choose your answers first. Explanations appear only after submission.

1. What is the strongest reason to preserve fictional source provenance after normalization?

2. Why should event time and collection time remain separate?

3. A fictional source introduces a new undocumented result value. What is the safest handling?

4. What is the main risk of normalization?

5. A fictional source reconnects after a blind period. Which state is strongest initially?

6. Which fictional collection metric is most meaningful?

7. Which portfolio approach is safest?

Portfolio Prompt

Portfolio Prompt

Create a fully fictional Collection and Normalization Package for the Northbridge Student-Support Cooperative. Include mission, defender questions, source categories, source owners, source purpose, source scope, exclusions, source events, source schemas, schema versions, parser versions, source fields, field types, allowed values, local meanings, required fields, optional fields, collection paths, collector identifiers, event time, collection time, processing time, indexing time, alert time, clock state, ingestion delay, queue age, retries, replay, duplicate states, out-of-order states, blind periods, recovery states, parsing, parse failures, Unknown values, field mappings, canonical fields, normalized values, transformations, enrichment sources, derived fields, identity context, device context, service context, destination context, authorization context, ownership context, change context, mission context, storage, indexing, searchable fields, access roles, retention, deletion, privacy purpose, coverage, source-health states, Healthy behavior, Conditional behavior, Degraded behavior, Blind behavior, Conflicting behavior, Recovering behavior, expected-record tests, missing-required-field tests, missing-optional-field tests, schema-drift tests, duplicate tests, out-of-order tests, blind-period tests, recovery-replay tests, expected outcomes, observed outcomes, defects, corrective actions, validation gates, collection-completeness metrics, field-completeness metrics, semantic-quality metrics, freshness metrics, delay metrics, duplicate metrics, parse-failure metrics, coverage metrics, privacy metrics, normalization debt, owner matrix, change history, review triggers, rollback, residual risks, mapping retirement, source retirement, architecture diagram, leadership summary, reflection, and a statement that every organization, source, schema, field, record, timestamp, owner, decision, and outcome is invented.

Show the fictional source value and normalized value together when semantic differences matter.
Keep event time, collection time, processing time, indexing time, and alert time separate.
Preserve Unknown and missing-data states rather than forcing familiar categories.
Use source health, privacy, ownership, testing, metrics, review triggers, rollback, and retirement as part of the pipeline design.
Keep the entire artifact completely fictional, defensive, non-operational, privacy-safe, evidence-aware, maintainable, and suitable for a public learning portfolio.

Confidence / Readiness Reflection

Are You Ready for Correlation and Alert Rules?

Before moving to A6.3, rate your readiness from 1 to 5 for source purpose, schemas, parsing, field mapping, normalization, enrichment, provenance, timing, source health, duplicates, recovery, privacy, coverage, quality, ownership, and complete fictionalization.

I can explain how fictional evidence changes from source record to normalized SIEM record.
I can preserve source-specific meaning after mapping to shared fields.
I can separate event, collection, processing, indexing, and alert time.
I can handle Unknown values, missing fields, duplicates, out-of-order records, blind periods, and recovery replay.
I can make source health and coverage visible in normalized evidence.
I can use privacy, access, retention, deletion, ownership, and review triggers.
I can evaluate collection quality using completeness, semantic quality, freshness, duplicates, failures, coverage, and debt.
I can produce a safe fictional pipeline without copying real schemas, fields, records, or systems.
Record one fictional source, one source field, one normalized field, one semantic limitation, one timing risk, one source-health state, and one question you will carry into A6.3.

Key Takeaways

What You Should Remember

1.Collection and normalization shape every later fictional SIEM search, correlation, alert, dashboard, case, metric, and leadership decision.
2.Raw evidence, parsed fields, normalized fields, transformed values, enrichment, derived context, and analyst interpretation are different layers.
3.Provenance should preserve fictional source category, schema, parser, source field, source value, transformation, timing, owner, and limitation.
4.Event time, collection time, processing time, indexing time, and alert time should remain separate.
5.Normalization supports consistency but can hide semantic differences when source values are treated as identical.
6.Unknown values, missing fields, duplicates, out-of-order records, blind periods, and recovery replay require explicit behavior.
7.Healthy, Conditional, Degraded, Blind, Conflicting, and Recovering states should affect pipeline and analyst confidence.
8.Collection quality includes completeness, semantic quality, freshness, delay, duplicates, parse failures, coverage, privacy, and debt.
9.Purpose limitation, field minimization, access, retention, deletion, ownership, testing, review triggers, rollback, and retirement belong in the design.
10.Every CyberShield collection artifact must remain fully fictional, authorized, defensive, non-operational, privacy-safe, and incapable of exposing real systems or people.

Navigation

Continue Module A6

Next, learn how fictional SIEM correlation and alert rules connect normalized records through identity, device, service, destination, session, request, change, timing, counts, sequences, states, source health, alternatives, and missing-data behavior.