Skip to content

datalake_fdw: write path — INSERT to data files and a committed snapshot #2018

Description

@MisterRaindrop

Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is C.

Scope

  • write.c, the orchestration the format interface was written for: a writer per segment, rolling at the target size, Iceberg file naming under <location>/data/, a FileMeta per file (counts, sizes, per-column bounds and null counts) for commit.
  • AM insert callbacks (tuple_insert, multi_insert, finish_bulk_insert): slots into the batch builder, field ids from the Iceberg schema into WriterOptions.field_ids.
  • Segments write straight to object storage; the QD gathers every FileMeta and commits once at pre-commit. Abort anywhere leaves no data file (resource owner plus object delete) and no snapshot.
  • COPY ... FROM uses the same path.
  • With it: datalake_fdw: NUMERIC/DECIMAL in the Parquet format layer #1988 NUMERIC and datalake_fdw: timestamp columns in units other than microseconds #1990 timestamp units.

Out of scope

UPDATE/DELETE (E), partitioned tables (#1683 §2.3).

Depends on

A, B2, B4. Rolling, FileMeta and abort logic are testable with the local-file functions first.

Acceptance

  • A multi-segment INSERT ... SELECT yields files readable by pyarrow and Spark; manifest counts and bounds match the data.
  • Nothing visible before commit; abort leaves no files and no snapshot.
  • Row group and target file size honoured within one row group.
  • A large insert stays within gp_vmem_limit_per_query; refusal surfaces as ERRCODE_OUT_OF_MEMORY.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    datalakecontrib/datalake_fdw and contrib/datalake_agent: Iceberg lake tables

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions