AW-10865990051
IOblend for Data Engineers · Apache Spark · DataOps

Build production data pipelines without rebuilding the production plumbing.

IOblend gives data engineers a metadata-driven development and execution layer on Apache Spark. Build batch, CDC and streaming pipelines with visual composition, SQL and Python, then test, debug, version, validate and run them with production DataOps controls already attached.

Use low-code where it removes repetition. Use SQL and Python where the logic deserves code. IOblend is not a no-code replacement for engineering judgement.

SOURCE → IOBLEND ENGINE → TARGETtransform in flight, keep the controls with the flow
SourcesDB · API · files · eventsBatch, static, CDC and streaming inputs.
IOblendSQL · Python · SparkState, joins, quality, contracts, lineage, errors.
TargetsLakehouse · DB · API · appWrite governed results where the architecture needs them.
Designerbuild + test
EngineSpark execution
Playbookportable JSON
Runtimecustomer infra
The engineering loop

Design, validate and operate the same pipeline model from desktop to cluster.

The useful unit is not a drag-and-drop diagram. It is a versioned production pipeline with explicit dependencies, executable logic, testable components, runtime configuration and DataOps behaviour that survives deployment.

01 · DesignCompose the graphSources, transforms, sinks, dependencies and run order in the Designer.
02 · CodeAdd SQL + PythonUse code inside the pipeline where transformations or custom logic require it.
03 · TestRun components locallyInspect intermediate data and validate logic before treating the flow as deployable.
04 · ControlAttach DataOpsSchema, quality, contracts, lineage, state and exception behaviour move with the pipeline.
05 · VersionStore the playbookPipeline metadata is JSON and can live in a normal source-code repository.
06 · RunExecute on SparkMove from local engine execution to customer-controlled enterprise Spark infrastructure.
The pipeline is metadata, not a screenshot

Portable JSON playbooks separate pipeline intent from runtime infrastructure.

IOblend stores the pipeline as metadata: components, dependencies, configuration, transformation behaviour and operational controls. That makes the graph portable, versionable and reusable instead of binding the engineering logic to one visual session or one target platform.

Illustrative playbook metadata

Component
sourcetransformsinkAI step
Dependency
run orderparentsconditionstate
DataOps
schemaqualitylineageerrors
Runtime
pathsparamssecrets refengine
{ "pipeline": "customer_cdc", "execution": "spark", "logic": "portable", "environment": "parameterised" }
Repository-friendlyJSON metadata can be stored in Git or another code repository alongside the rest of the engineering estate.
Diffable versionsPipeline versions can be compared as the graph, parameters and logic evolve.
Infrastructure-independent logicKeep the repeatable data logic separate from connection details and runtime-specific values.
Reusable playbooksTurn similar integration families into parameterised patterns instead of copied pipelines.
Batch, CDC and streaming in one production model

Change the arrival pattern without rebuilding the entire engineering stack.

IOblend uses a Kappa-style architecture so batch, database changes and event streams can participate in the same transformation and DataOps model. That matters when a pipeline needs a static reference table, live events and a CDC feed in the same logical flow.

Three arrival patterns, one transform graph

Batch
CDC
Streaming
IOblend engineshared transforms · state · quality · lineage
Static + streaming joinsEnrich live events with lookup or reference datasets inside the same pipeline.
In-flight transformsApply business logic before the data reaches its target rather than staging first by default.
Common DataOpsUse the same quality, schema and lineage controls regardless of arrival mode.
One operational modelReduce the number of frameworks engineers have to support just because latency requirements differ.
Stateful streaming without hand-building the state manager

Window, deduplicate, upsert and handle late events as part of the pipeline.

Real-time engineering gets difficult when the flow becomes stateful. IOblend supports production patterns such as windowed processing, chained aggregations, deduplication, SCD logic, real-time upserts and late-arriving data while keeping the state behaviour visible in the same pipeline model.

Illustrative event-time window

LATE EVENT
earlier event timelater event time
DeduplicationMaintain deterministic downstream state when retries or source behaviour produce duplicate events.
Idempotent outcomesDesign repeated processing so the target state does not keep changing because an event was seen twice.
Windowed stateAggregate or correlate events over defined windows instead of treating every event as stateless.
Late data handlingDeal with out-of-order arrivals without hiding the problem inside custom recovery code.
Streaming deduplication deep dive →
CDC as an engineering primitive

Capture database changes, transform them in flight and write the result where it belongs.

IOblend supports hybrid CDC approaches including log-based, trigger-based and query-based patterns. Engineers can apply transforms, quality logic and schema controls to the change stream before upserting into a lakehouse, database, application or synchronised target.

Illustrative change sequence

INScustomer_id 48201lsn:8812
UPDstatus ACTIVElsn:8813
DELobsolete rowlsn:8814
UPDaddress normalisedlsn:8815
INScustomer_id 48202lsn:8816
Hybrid captureSelect the CDC mechanism appropriate to the source and operational constraints.
Transform before sinkNormalise, enrich or validate change events before they become target state.
Schema-aware changesTrack structural evolution instead of treating the database change stream as schema-free.
Synchronisation patternsUse CDC for lakehouse ingestion, migration tailing, system sync and operational replication.
CDC architecture deep dive →
Data contracts that execute

Treat schema evolution as a controlled change, not a 02:00 production surprise.

IOblend can generate and version schemas, validate incoming records against expected contracts and isolate records that violate the contract. The goal is to make structural change explicit while allowing healthy records to continue where the pipeline policy permits it.

Illustrative schema diff

v17
customer_idLONG
postal_codeSTRING
statusSTRING
v18
customer_idLONG
postal_codeINT?
statusSTRING
Contract violation → quarantine affected records → preserve valid flow
Schema generation + versioningKeep an explicit history of structural expectations as upstream systems evolve.
Automatic validationCheck batches and streams before malformed structure becomes target state.
Record isolationRoute invalid records to an error path instead of forcing every quality problem to crash the whole pipeline.
Trace the impactUse lineage and metadata to identify which record changed, where it came from and what logic touched it.
CI/CD starts before the commit

Test-as-you-build, then treat the playbook like software.

IOblend validates the pipeline during construction. Components can be run and inspected before deployment, invalid pipeline logic is blocked, and SQL or Python syntax problems are surfaced before execution. Pipeline versions are stored and the JSON playbook can participate in normal repository and deployment practices.

Integrated validation path

Componentrun
SQL / Pythonsyntax
Graphlogic
Versiondiff
Deployruntime
Invalid logic is blocked before the pipeline is treated as executable.
REPL-like inspectionRun pipeline components locally and inspect intermediate datasets while developing the flow.
Visual debuggingTrace behaviour through the graph rather than reconstructing the entire pipeline from distributed logs.
Version comparisonCompare pipeline versions as metadata, configuration and code evolve.
Repository integrationStore JSON playbooks in the engineering repository and promote them through your own delivery process.
Visual Spark debugging deep dive →
Debug the record, not just the job

Carry lineage and exception context through the transformation path.

Pipeline-level logs tell you that a job failed. Record-level lineage helps answer the harder questions: which record changed, which component changed it, what source it came from and why it was routed to an exception path.

Illustrative record trace

Sourcecrm.customer
Mapnormalise
Joinaccount
Validatecontract
Sinkcustomer_360
Exception path retains record + transformation + source context.
Record-level tagsCarry lineage context as data moves through the pipeline rather than reconstructing it after the fact.
Error isolationSeparate problematic records from healthy flow and preserve the information needed to investigate.
Operational metadataExpose pipeline state and execution context without bolting on a second observability model for basic questions.
Faster replay decisionsKnow what failed and where before deciding whether to fix, replay, amend or reject.
Develop locally. Execute on customer-controlled Spark.

The deployment boundary should not change the pipeline logic.

Developer Edition installs the Designer and a local IOblend Engine with a local Spark environment. Enterprise Edition uses a remote IOblend Engine packaged for customer cloud or on-premises Spark infrastructure. The Designer can connect to local or remote engines for development and testing.

Promotion path

Developer EditionDesigner + local EngineBuild and execute pipelines on the developer machine while connecting to accessible cloud or on-prem sources and sinks.
Enterprise EditionRemote Engine on SparkExecute generated run files in the customer's Spark environment.
Airflowschedule
Enterprise schedulertrigger
Repositoryversion JSON
Customer infrastructureDeploy in cloud, on-premises or hybrid environments instead of sending production data to a mandatory IOblend SaaS plane.
Environment parametersChange runtime configuration without rewriting the business and transformation logic.
External schedulingGenerated run files can participate in existing scheduling and orchestration practices such as Airflow.
Same engineering modelUse the same Designer and pipeline representation from desktop development into enterprise execution.
AI inside ETL, not beside it

Use model or agent logic as a controlled pipeline component when the dataflow needs it.

IOblend can embed Python and Agentic AI logic inside the ETL path for tasks such as extracting fields from documents, classifying content or validating unstructured inputs. The output can then pass through the same structured quality, quarantine and lineage model as the rest of the pipeline.

Controlled AI step inside the graph

Document / textunstructured input
Python / Agentextract or classify
Validated recordstructured output
Keep AI scopedUse the model for the task it is good at, then return the output to explicit pipeline logic.
Validate the resultCheck extracted fields or classifications against required structure before downstream use.
Quarantine uncertaintySeparate low-confidence or invalid outputs instead of allowing them silently into target data.
Join back to enterprise contextCombine approved unstructured extraction with existing customer, asset, order or operational records.
Agentic AI pipeline architecture →
Use the product, not just the architecture diagram

Start locally, follow the pipeline tutorials, then move into production behaviour.

The Developer Edition includes the Designer, local Engine and local Spark environment. The documentation already provides a practical sequence from installation and run parameters through static pipelines, streaming, event management and JDBC sinks.

Engineering ecosystem context

IOblend should fit your engineering practices, not pretend the rest of the ecosystem does not exist.

Work with your existing stack and augment where you have gaps and performance issues. Use your preferred schedulers, ontology and observability tools as needed.

IOblend's role: provide the development, metadata, DataOps and execution layer around the pipeline. It does not replace Spark itself, your scheduler, your source repository or engineering judgement.
Data engineering FAQ

What IOblend actually changes in the engineering workflow.

These answers are written for engineers evaluating the product, not for a generic platform comparison.

Apache SparkCDCstreamingDataOpsdata contractslineage
What is IOblend for a data engineer?

IOblend is a metadata-driven data integration and DataOps development layer on Apache Spark. Engineers use the Designer, SQL and Python to build batch, CDC and streaming pipelines, then execute those pipelines with built-in quality, schema, lineage, state and exception controls.

Does IOblend generate Apache Spark pipelines?

Yes. The IOblend Engine turns the playbook definition into distributed Spark execution so engineers do not have to hand-build the surrounding Spark application for every production pipeline.

Is IOblend no-code?

No. IOblend supports low-code visual composition to remove repetitive engineering, while SQL and Python remain available for custom transformation and processing logic.

Can IOblend handle batch and streaming in the same pipeline?

Yes. Static, batch, CDC and streaming inputs can participate in the same pipeline model, including joins between live events and reference data.

How does IOblend handle Change Data Capture?

IOblend supports hybrid CDC patterns including log, trigger and query approaches, then allows engineers to transform and validate change events before writing target state.

How does IOblend handle duplicate events and idempotency?

Stateful deduplication and upsert patterns can be used so retries or duplicate events do not create repeated target outcomes.

How are schema drift and data contracts handled?

IOblend can generate and version schemas, validate data against defined contracts and isolate invalid records while valid data continues where the pipeline policy allows it.

What does test-as-you-build mean?

Pipeline components can be executed and inspected during construction. IOblend also validates pipeline logic and surfaces SQL or Python syntax errors before the pipeline is treated as executable.

Can IOblend pipelines be version controlled?

Yes. Pipeline definitions are stored as JSON metadata files and can be stored in a normal source-code repository. IOblend also stores pipeline versions for comparison.

How does visual debugging work?

Engineers can inspect components and intermediate data through the development environment, which reduces the need to infer the entire pipeline state only from distributed runtime logs.

What is record-level lineage in IOblend?

IOblend attaches lineage context at record level as data moves through the flow, allowing engineers to trace source, transformation and exception context for individual records.

Where does IOblend run?

Developer Edition runs the Designer and local Engine on the developer machine. Enterprise Edition can run a remote Engine on customer-controlled cloud, on-premises or hybrid Spark infrastructure.

Can I use Airflow with IOblend?

Yes. Enterprise run files can be scheduled by external scheduling software such as Apache Airflow, allowing IOblend to participate in existing orchestration practices.

Can I try IOblend before an enterprise deployment?

Yes. Developer Edition is available for local development and the documentation includes installation, first-pipeline, streaming and JDBC tutorials.

Bring a real pipeline

Start with the source, the state problem and the production controls you are tired of rebuilding.

Use Developer Edition to build locally, or bring us a representative CDC, streaming, migration or batch pipeline and we can work through how the playbook, Engine and DataOps model would apply.

Scroll to Top