Subsurface Analytics

Iterative empirical/deterministic and statistical (ML) tests performed on subsurface measurement data -- lithology classification, core-log depth QC, and water-saturation calibration. Both approaches are applied to the same wells to allow direct comparison of results. Each card can be re-run as underlying data evolves.

Core-Log Depth Misalignment QC

Flags core samples whose density-derived porosity disagrees with lab-measured porosity, at both the individual-sample and physical-core-run level -- a QC signal for bad depth ties.

wells
1326
groups
5189
clean groups
4394
matched pairs
212446

Last run: 2026-07-29 13:07

Open

Core-Log Shift-Search (Depth Correction Estimate)

Companion to the depth-misalignment QC test: for each core run identified as mis-aligned, an estimate of the depth correction is determined by 'sliding' the core samples against the log and finding the best-fit alignment.

skipped
0
candidates
360
inconclusive
71
confirmed shift
222

Last run: 2026-07-23 03:47

Open

Core-Log Porosity Crossplot (by Formation)

Compares real lab-measured core porosity against the project's own computed log PHI curve, at the same depth, per sample -- not averaged, so real scatter and bias stay visible. Formation is resolved via the covering stratigraphic interval (core samples don't carry a usable formation field of their own); pick up to 10 formations at a time directly on the chart, or use Display Top 10 for the formations with the most core coverage. The biggest formations are randomly capped to a fixed sample count so the chart stays legible -- see the summary stats for how many of each formation's samples are actually shown.

n wells
1211
n points
131579
n formations
498
mean residual
0.17

Last run: 2026-08-09 18:57

Open

Core PHI vs. Log PHI, by Supplier

Compares real lab-measured core porosity against EACH available porosity supplier (the project's own computed PHI curve, plus any third-party CPI vendor curve), one supplier at a time, as a binned density heatmap -- the companion Core-Log Porosity Crossplot card groups the SAME underlying matches by formation instead of by supplier, and shows raw points rather than a heatmap. Pick a supplier chip to see its own agreement pattern against core; the summary table below ranks all suppliers by accuracy.

n wells
1212
n points
180721
n suppliers
5

Last run: 2026-08-09 20:19

Open

Empirical Lithology Crossplot (PEF/RHOB/DT) -- Test Wellbores

Empirical lithology classification check: despiked PEF/RHOB/DT medians per lithology-labeled interval, gated to GR-clean samples only (Clavier Vshale <= 0.15), for 6 hand-picked test wellbores, against published petrophysical matrix endpoints. Scoped to a small test set while the method is developed.

wells
6
vsh cutoff
0.15
class summaries
16
deviation flags
7

Last run: 2026-07-24 03:03

Open

PE/Uma Matrix Crossplot -- Representative Sample (WellStrat Pure Intervals)

PE, the photoelectric factor measured by the density tool (curve mnemonic PEF), is a strong, largely porosity-independent indicator of a rock's mineral composition. Uma is PE's zero-porosity matrix equivalent -- what the photoelectric response would read for a pure sample of that mineral with no pore space.

wells
412
vsh cutoff
0.15
class summaries
358
deviation flags
43

Last run: 2026-08-04 17:56

Open

Empirical Lithology Classifier (Phase 1) -- Full Database

Deterministic lithology classifier with ZERO fitted parameters. GR is a prerequisite curve, via a Clavier Vshale gate. A tiered rule then splits clean samples into two checks, then applies optional refinements:

wells
2567
n total
22900975
coverage pct
37.5
clean vsh max
0.15

Last run: 2026-08-02 17:46

Open

Statistical Lithology Classifier (Phase 2) -- XGBoost

Lithology classifier based on an XGBoost model -- an ensemble of 200 stacked decision trees, where each new tree is trained specifically to correct errors from the trees before it ("gradient boosting"). It learns its patterns from real examples, rather than applying fixed rules, and is comparable to Phase 1's rule-based approach, not combined with it. It predicts 10 rock types: Shale, Sandstone, Limestone, Halite, Chalk, Anhydrite, Marl, Dolomite, Tuff, and Coal.

n total
6250594
test wells
521
train wells
2103
training rows
13369092

Last run: 2026-08-02 15:43

Open

Statistical Lithology Classifier (Phase 3) -- Purity-Weighted XGBoost

Same XGBoost approach as Phase 2, with one change: Phase 2 treats every WellStrat-picked interval as 100% pure the rock type it's labelled -- but checking the database shows only 30.6% of the intervals behind these classes actually are 100% pure; roughly a quarter are below 75% pure, real interbedding and mixing, not noise. Phase 3 weights each training sample by its own recorded purity (WellStrat's litho1_perc), so a 75%-pure Sandstone interval counts for 0.75x a training example, not the same as a genuinely 100%-pure one. A hard cutoff (only train on 100%-pure intervals) was considered and rejected -- it would leave Coal, already the thinnest class, with just 14 rows database-wide, below the minimum needed to train on it at all. Weighting lets impure intervals still contribute, in proportion to how trustworthy their label actually is.

n total
5376253
row stride
3
test wells
473
train wells
1905

Last run: 2026-08-09 14:27

Open

Phase 4: SANDCOUNT Binary Sand Classifier

A different question from Phase 2/3: not "what rock type is this", but simply "is this Sand or not" -- trained on nsta.lithology_interval (the Lloyd's Register SANDCOUNT product), not WellStrat. SANDCOUNT barely labels the fine fraction (Shale/Siltstone/Mudstone are almost never picked), so this card deliberately treats ANY depth not picked as Sand/Sandstone as a confirmed Not-Sand example -- see the ground-truth note below. Reports precision/recall at several probability thresholds rather than one fixed cutoff, since recall on Sand (missed potential reservoir) matters more here than a single accuracy number.

pr auc
0.7229
n total
1988371
roc auc
0.8973
test wells
103

Last run: 2026-08-02 15:25

Open

Rw Calibration -- Pickett Plot

Stage 1 of the Archie/Sw hydrocarbon-detection pipeline: fits formation water resistivity (Rw) per reservoir formation via a Pickett-plot lower-envelope regression, across 10 target reservoir formations. Rw is a required input to Archie's equation for computing water saturation (Sw) -- not yet built, this is calibration only.

formations ok
6
formations fitted
10

Last run: 2026-08-02 10:25

Open

Rw Formation Grouping Check

Does grouping by formation actually matter for Rw, or would one shared number work almost as well? Fits Rw separately for every single well (not pooled with other wells), then checks whether wells in the same formation land close together and clearly apart from other formations -- or whether the numbers look about equally scattered no matter which formation a well belongs to.

well formation fits
753
formations with fits
10
between formation std
1.5789
avg within formation std
5.8036

Last run: 2026-07-31 10:02

Open

Rw Calibration (All Formations, Data-Discovered)

Same Pickett-plot Rw fit as the main Rw Calibration card, but not limited to a hand-picked list of 10 formations -- every formation with enough qualifying wells is found directly from the data. Built after a real gap was found this session: well 9/09a- 11's two correctly-recorded gas shows (71m and 97m thick) sit in Beryl Formation, which was never in the original 10 and so never entered Rw/Sw at all. Writes its own separate results -- the original 10-formation card is completely unaffected by this one.

min wells
10
formations ok
107
formations fitted
196
formations discovered
224

Last run: 2026-08-02 10:33

Open

Stage 2: Sw Validation (All Formations, Data-Discovered)

Same Sw validation as the main Stage 2 card, but over formations discovered directly from the data instead of the fixed 10 -- so real, known pay sitting outside that original list (like well 9/09a- 11's Beryl Formation gas shows) now actually gets checked. Requires the 'Rw Calibration (All Formations)' card to have been run first -- this reads the Rw values that card saves.

min wells
10
total samples
3027276
qualifying wells
2138
formations covered
208

Last run: 2026-08-01 18:49

Open

Rw Calibration (Per Sand Interval)

A different unit of calibration entirely: instead of pooling wells sharing a formation name, or fitting one number per well, Rw is fitted separately for every individual water-bearing sand interval that's thick enough to support it. A single sand interval is one genetically coherent deposit -- assuming uniform water chemistry within one is far more defensible than assuming it across a whole well (which can cross several unrelated sand bodies) or a whole named formation (which pools sand bodies from many wells that may never have been in contact). Every interval's fit still carries its own formation, so a real distribution of Rw builds up per formation from genuine evidence.

interval fits
6600
wells processed
1782
formations with fits
462

Last run: 2026-08-02 17:00

Open

Sw by Stratigraphy (Wells South-North)

Visualizes the range of Sw seen in sand bodies across wells, organized by stratigraphic unit (Group/Formation/Member level) and by well position South to North. Each point is one sand interval's median Sw, computed with that SAME interval's own fitted Rw from the Rw Calibration (Per Sand Interval) card -- not a formation-level average -- so the granularity matches how Rw was actually calibrated. Not restricted to a fixed formation list: any stratigraphic unit with interval-level Rw evidence can be picked, up to 10 at a time, directly on the chart.

n wells
1847
n strat units
823
n interval points
7467

Last run: 2026-08-02 05:15

Open

Curve-Shape Clustering (Overlooked Pay)

Unsupervised clustering of raw wireline curve shape (GR, RHOB, NPHI, DT, true resistivity, Vsh, porosity -- NOT Sw or any hydrocarbon flag) across every Phase-4-predicted-Sand interval project-wide, then cross-checked against WellStrat's own hydrocarbon flag and computed Sw for enrichment. The point: a model trained directly on the hydrocarbon flag just re-learns historical picking bias -- clustering with zero reference to that flag can surface curve-shape groups that look like known pay but weren't flagged, i.e. candidate overlooked pay.

k
5
n wells
1695
silhouette
0.2271
n interval points
7027

Last run: 2026-08-02 05:20

Open

Sw Predictor (Direct Supervised, Merkle Aquila Model 1)

Trains an XGBoost regressor to predict Sw directly from wireline curve shape -- NOT the hydrocarbon flag. Checked against the real label data first: only 12 of ~105k strat_interval rows have an explicit 'checked, not hydrocarbon' code, so a classifier trained on the flag would just re-learn WellStrat's own historical picking pattern. Predicting the same continuous Sw the Archie physics chain already computes sidesteps that -- no hydrocarbon flag anywhere in the loop -- and, once trusted, could eventually estimate Sw wherever curves exist but no Rw fit does.

r2
0.6145
rmse
0.1561
n total
1122076
test wells
1082

Last run: 2026-08-04 06:47

Open

Coal Identification & Rank Classifier (Crain)

Deterministic Coal analysis, sourced from Crain (2010) 'Unicorns in the Garden of Good and Evil, Part 2 - Coal', CSPG Reservoir. Independently re-identifies coal via a 5-curve threshold-flag rule (NPHI high, RHOB low, DT high, resistivity high, GR clean -- 3-of-5 triggered), WITHOUT trusting the SANDCOUNT 'Coal' label to decide where coal is, then measures how often the two agree. Where they agree, nearest-endpoint matches a coal RANK (Anthracite/Bituminous/Lignite/Peat) and decomposes RHOB into Ash/Fixed-Carbon/Moisture/Volatile fractions. Zero fitted parameters, no training data -- same philosophy as the Phase 1 classifier. Built to directly test the open question flagged in the PEF/RHOB crossplot card: SANDCOUNT Coal picks previously looked inconsistent with coal's expected log signature in every well examined.

wells
197
untestable
2242
agreement pct
38.7
threshold agree
3476

Last run: 2026-08-02 08:39

Open

Coal Method Comparison -- Threshold-Flag vs. Single Centroid vs. 4-Rank

Head-to-head comparison of the three approaches discussed for adding Coal to the empirical classifier: (1) Crain's threshold-flag rule alone, (2) Coal as ONE mineral-style endpoint (Bituminous, same treatment as Sandstone/Limestone/Halite/Anhydrite), and (3) the full 4-rank mixing model (the 'Coal Identification & Rank Classifier' card). Run on the FULL depth range of every well with a SANDCOUNT Coal pick, not just inside the picks -- gives real precision, not recall alone.

wells
197

Last run: 2026-08-02 08:49

Open