mirror of
https://github.com/Jeuners/ECC.git
synced 2026-09-16 01:56:10 +02:00
* fix: make the installer runtime pass strict supply-chain vetting
Remediate the four enterprise supply-chain vetting blockers from
affaan-m/ECC#2502 so the installer runtime (package.json + manifests +
scripts/lib/**) passes strict exact-pin evidence policy:
1. Remove the package.json `postinstall` lifecycle script (it only echoed a
post-install banner) and move that banner to an explicit opt-in
`npm run welcome` command. No install-time lifecycle script remains.
2. Exact-pin every dependency in package.json (dependencies + devDependencies)
to the versions already resolved in package-lock.json; no ^/~ ranges.
3. Replace non-ASCII characters on the installer runtime script/config surface:
em-dashes (U+2014) in scripts/lib/{path-safety,install-executor,
install/link-rewrite}.js comments and the two "Itô" (U+00F4) occurrences in
manifests/{install-components,install-modules}.json descriptions become
ASCII, so strict-surface Unicode scanners are clean.
4. Drop the bare `require("ajv")` from scripts/lib/install-state.js; the file
already carries a complete hand-rolled validator enforcing the same
schemas/install-state.schema.json (ecc.install.v1) constraints, so the
installer closure is dependency-free (zero non-builtin bare requires).
Refs affaan-m/ECC#2502
* fix: avoid unpinned welcome invocations
Signed-off-by: Samar Tomar <samar_tomar@hotmail.com>
* fix: validate translated skill frontmatter
Signed-off-by: Samar Tomar <samar_tomar@hotmail.com>
* fix: repair skill frontmatter YAML
Signed-off-by: Samar Tomar <samar_tomar@hotmail.com>
* fix: add MIT license to core skill manifests; pin verification-loop tsc invocation
* fix: preserve tsc/pyright exit status in verification-loop type-check (set -o pipefail)
* chore(deps): sync lockfiles with exact-pinned package.json
Regenerate package-lock.json and yarn.lock so the pinned dependency
specs are reflected in both lockfiles. npm ci and Yarn's --immutable
install now pass the sync check. The resolution tree is unchanged
(231 yarn resolutions, byte-identical set; zero npm transitive drift);
only the root descriptor strings move from ranges to the versions
already resolved in the committed lockfiles.
Addresses the Codex P1 on #2503.
---------
Signed-off-by: Samar Tomar <samar_tomar@hotmail.com>
Co-authored-by: Samarjeet Singh Tomar <samartomar@gmail.com>
74 lines
2.6 KiB
Markdown
74 lines
2.6 KiB
Markdown
---
|
|
name: data-throughput-accelerator
|
|
description: Use when large data ingestion, backfill, export, ETL, warehouse loading, manifest catch-up, or table synchronization needs to become much faster while preserving data correctness.
|
|
license: MIT
|
|
metadata:
|
|
origin: ECC
|
|
tools: Read, Write, Edit, Bash, Grep, Glob
|
|
---
|
|
|
|
# Data Throughput Accelerator
|
|
|
|
Use this skill when the bottleneck is moving, transforming, or saving lots of
|
|
data. The goal is not just speed. The goal is faster correct data landing in the
|
|
right place with proof.
|
|
|
|
## First Distinction
|
|
|
|
Separate these before optimizing:
|
|
|
|
- source extraction speed;
|
|
- network transfer speed;
|
|
- warehouse/load speed;
|
|
- transform speed;
|
|
- serving-table freshness;
|
|
- live tail growth while the job runs.
|
|
|
|
A pipeline can be "fast" and still appear behind if new data arrives faster
|
|
than the final catch-up window.
|
|
|
|
## Fast Path Heuristics
|
|
|
|
- Move compute to where the data already is.
|
|
- Prefer warehouse-native scans, joins, and appends for large landed files.
|
|
- Use manifests or checkpoints so completed files/partitions are skipped.
|
|
- Use partitioning and clustering that match the read and append pattern.
|
|
- Batch small files, requests, and writes.
|
|
- Make writes idempotent through unique keys, manifests, or replaceable staging.
|
|
- Keep raw, derived, and serving tables separately accountable.
|
|
|
|
## Workflow
|
|
|
|
1. Read the current source, target, and manifest contracts.
|
|
2. Measure backlog: external files, manifest rows, raw rows, derived rows,
|
|
min/max timestamps, and unprocessed counts.
|
|
3. Run a safe catch-up or sample benchmark.
|
|
4. Compare variants: batch size, worker count, warehouse SQL, file grouping,
|
|
staging shape, and manifest update method.
|
|
5. Promote only the fastest path that keeps counts and timestamps coherent.
|
|
6. Codify the path as a CLI, scheduled job, workflow, or runbook.
|
|
7. Rerun final accounting after the codified path executes.
|
|
|
|
## Accounting Output
|
|
|
|
Use a hard accounting block:
|
|
|
|
```text
|
|
Data throughput result:
|
|
- Source files discovered: 294
|
|
- Files processed this run: 294
|
|
- Raw rows added: 9,683,598
|
|
- Derived rows added: 8,917,585
|
|
- Remaining tail: 24 files at readback time
|
|
- Runtime: 38.7s
|
|
- Correctness gate: manifest counts and table max timestamps match
|
|
```
|
|
|
|
## Guardrails
|
|
|
|
- Do not delete raw data to make a metric look better.
|
|
- Do not skip failed files silently.
|
|
- Do not mix historical backfill status with live-tail freshness.
|
|
- Do not call a pipeline complete until the target tables and manifest agree.
|
|
- For finance, healthcare, regulated, or customer-impacting data, preserve
|
|
replay evidence and approval gates.
|