# Apache ORC C++ Reader — Unbounded Allocation & Integer Overflow Report ## Overview Four (4) security findings in the Apache ORC C++ core reader (`orc::createReader()` / `RowReader::next()`) discovered via 6.5-hour AFL++ fuzzing campaign and manual PoC development. **Target:** Apache ORC C++ v2.3.1 (commit `a6f12fd`, 2026-07-17) **Toolchain:** clang-18 + ASan + UBSan + trace-pc-guard **Fuzzer:** 3 AFL++ workers, 1M+ execs, 23.81% coverage **Reproducer:** Any program calling `orc::createReader()` + `RowReader::next()` on the crafted `.orc` files --- ## Finding 1: DictionaryLoader — Integer Wraparound (CWE-190) **File:** `c++/src/DictionaryLoader.cc:67` **PoC:** `dictionary_wraparound_poc.orc` `uint32_t dictSize = encoding.dictionary_size()` — no upper-bound validation. When `dictSize = 0xFFFFFFFF`, `dictSize + 1` wraps to `0` in uint32 arithmetic. `dictionaryOffset.resize(0)` allocates a zero-capacity buffer, and `lengthDecoder->next()` writes ~4.29 billion `int64_t` values past it → heap-buffer-overflow. **Confirmation:** ASan heap-buffer-overflow WRITE (standalone harness + real ORC file) **Prior Art:** ❌ EnigmaConsultant/orc-dictionary-overflow (same root cause, included for completeness) **Severity:** High (OOB write from crafted column encoding metadata) --- ## Finding 2: DictionaryLoader — Unbounded Allocation (CWE-789) **File:** `c++/src/DictionaryLoader.cc:67` **PoC:** `dictionary_unbounded_alloc_poc.orc` Same missing upper bound on `dictSize`, but with a large positive value (e.g., `0xfffff9cd = 4294965709`) instead of the exact wraparound boundary. This causes `resize(4294965710)` → ~34GB allocation → ASan OOM / allocation-size-too-big crash. **Confirmation:** ASan allocation-size-too-big (fuzz crash from 6.5h campaign) **Novelty:** ✅ Not covered by EnigmaConsultant (whose PoC only tests the exact 0xFFFFFFFF wraparound) **Severity:** Medium-High (OOM / resource exhaustion via uncontrolled allocation) --- ## Finding 3: ColumnReader — Unbounded Allocation (CWE-789) **File:** `c++/src/ColumnReader.cc:743` (StringVectorBatch / StringDirectColumnReader) **PoCs:** `string_vectorbatch_resize_poc.orc`, `string_direct_reader_resize_poc.orc` Two crash files from the fuzzer that trigger `DataBuffer::resize(huge)` via: - `StringVectorBatch::resize()` called from `StringDictionaryColumnReader::next()` (745B crash) - `StringDirectColumnReader::next()` direct blob resize (373KB crash) Both paths decode untrusted stream lengths and allocate without upper-bound validation. **Confirmation:** ASan heap-buffer-overflow / allocation-size-too-big (fuzz-confirmed) **Novelty:** ✅ Novel — separate code path from DictionaryLoader **Severity:** Medium (OOB write / OOM) --- ## Finding 4: StripeStream — Integer Overflow → OOB Read (CWE-190 → CWE-125) **File:** `c++/src/StripeStream.cc:94` **PoC:** `stripestream_offset_overflow_poc.orc` **Confirmation:** Manually-confirmed (protobuf-valid PoC) ### Root Cause ```cpp // StripeStream.cc:82-126 uint64_t offset = stripeStart_; uint64_t dataEnd = stripeInfo_.offset() + stripeInfo_.index_length() + stripeInfo_.data_length(); for (int i = 0; i < footer_.streams_size(); ++i) { const proto::Stream& stream = footer_.streams(i); if (stream.has_kind() && stream.kind() == kind && stream.column() == columnId) { if (offset + stream.length() > dataEnd) { /* overflow-unsafe addition */ ... } ... } offset += stream.length(); /* cumulative, overflow-unsafe */ } ``` `offset += stream.length()` is applied to ALL streams regardless of whether they match the search. If a non-matching stream has `length = UINT64_MAX`, the cumulative offset wraps. A subsequent matching stream then bypasses the `offset + streamLength > dataEnd` check because the wrapped value is small. ### ASan Output (from PoC) ``` READ of size 200 at 0x... (1 byte before buffer) UBSan: addition of unsigned offset to 0x... overflowed to 0x... Stack: MemoryInputStream::read → SeekableFileInputStream::Next → RleDecoderV2::readByte → StringDirectColumnReader::next ``` ### Impact Heap-buffer-overflow **READ** (not WRITE). The wrapped offset causes the reader to read from memory before the allocated buffer, leaking or crashing. **Novelty:** ✅ Novel — different mechanism from DictionaryLoader, manually confirmed **Severity:** Medium (OOB read, potential info leak / crash) --- ## Fix Recommendations 1. **DictionaryLoader.cc:67** — Add upper-bound check on `dictionary_size`: ```cpp if (dictSize > MAX_DICT_ENTRIES) throw ParseError("..."); dictionary->dictionaryOffset.resize(static_cast(dictSize) + 1); ``` 2. **ColumnReader.cc** — Add upper-bound check on decoded lengths before `resize()`: ```cpp if (totalLength > MAX_BLOB_SIZE) throw ParseError("..."); byteBatch.blob.resize(totalLength); ``` 3. **StripeStream.cc:94** — Use overflow-safe addition for cumulative offset and data checks: ```cpp if (addOverflow(offset, stream.length(), &offset)) throw ParseError("..."); ``` --- ## Timeline | Date | Event | |------|-------| | 2026-07-19 | Fuzzing campaign starts (3 AFL++ workers) | | 2026-07-19 | 7 crashes captured, all traced to missing upper bounds | | 2026-07-19 | StripeStream overflow manually confirmed with protobuf-valid PoC | | 2026-07-19 | Report finalized | ## Credits - @drogba771 — Discovery, fuzzing, PoC, analysis