-
Notifications
You must be signed in to change notification settings - Fork 0
PRIK Vision
What PRIK is being built toward: a semantic multilanguage wrapper and interoperability runtime.
Status: design. This page states goals and reasoning. It is not a statement of current capability, and nothing here is a support claim.
Implemented today: the semantic model both frontends converge on, editable
.pyicontracts, API projection onto native calls, ownership and lifetime policy, and the Fortran and direct-C backends. See PRIK Architecture for the pipeline that exists and the feature matrix for what it supports.Design direction, not built: the coercion engine and registry, the runtime validation engine and contracts, additional backend adapters, the plugin ecosystem, and API style projection.
The goal of this project is to create a modern interoperability framework capable of wrapping and connecting libraries written in multiple native languages through a unified semantic API layer.
The system should:
- wrap native libraries from:
- Fortran
- C
- C++
- Rust
- CUDA
- and potentially more languages later
- expose clean Python APIs
- support semantic interoperability between different native runtimes
- support automatic coercions and conversions
- support runtime constraints and validation contracts
- support zero-copy array interoperability when possible
- avoid compiler dependence whenever possible
- avoid forcing users to modify native code
- avoid the limitations of SWIG/f2py-style systems
- support wrapping libraries even when the source code is unavailable
The project is not merely a wrapper generator.
It is a:
- semantic wrapper compiler
- interoperability runtime
- runtime coercion engine
- runtime validation engine
- runtime validation contract system
- language-independent semantic API system
The most important architectural decision is:
The semantic API layer is the source of truth.
NOT:
- parser ASTs
- compiler internals
- ABI details
- native language syntax
The system separates:
| Concern | Responsibility |
|---|---|
| Semantic API |
.pyi-style interface layer |
| Runtime coercions | conversion registry and coercion graph |
| Runtime validation | constraint checks on adapted values |
| Validation contracts | reusable preconditions, postconditions, and invariants |
| Native ABI | backend adapters |
| Source parsing | optional helper |
This separation is the foundation of the whole architecture.
SWIG is:
- parser-centric
- macro-heavy
- difficult to debug
- weak for scientific arrays
- poor for runtime semantics
- poor for modern interoperability
It also becomes difficult to maintain when:
- ownership becomes complex
- NumPy arrays are involved
- GPU arrays are involved
- runtime conversions are needed
- API remapping becomes advanced
f2py is:
- Fortran-specific
- procedural
- compiler/build-centric
- not semantic-runtime oriented
- weak for object systems
- weak for heterogeneous runtimes
pybind11 is excellent for:
- clean bindings
- modern Python APIs
But:
- bindings are handwritten
- there is no semantic interoperability layer
- no runtime coercion model
- no language-independent abstraction
The architecture is composed of multiple layers.
Native libraries/sources
↓
Optional parser frontends
↓
Canonical semantic interface layer (.pyi)
↓
Semantic IR
↓
Runtime coercion engine
↓
Runtime validation engine
↓
Runtime validation contracts
↓
Backend adapters
↓
Generated CPython extension
↓
Python API
The validation engine enforces concrete checks. Validation contracts describe when those checks run, what they guarantee, and how failures are reported.
The .pyi-style interface file is the central abstraction.
It defines:
- semantic APIs
- classes
- functions
- methods
- semantic types
- allowed coercions
- constraints
- validation contracts
- ownership semantics
- API remapping
This layer is:
- language-independent
- parser-independent
- editable
- human-readable
- stable
The native parser is NOT the source of truth.
The .pyi interface file is.
Suppose native Fortran code contains a procedural matrix API:
module sparse_mod
type :: sparse_matrix
end type
contains
subroutine create_sparse(A, nrows, ncols)
type(sparse_matrix), intent(out) :: A
integer, intent(in) :: nrows, ncols
end subroutine
subroutine sparse_multiply(A, x, y)
type(sparse_matrix), intent(in) :: A
real(8), intent(in) :: x(:)
real(8), intent(out) :: y(:)
end subroutine
end moduleThe semantic interface may expose a Pythonic object model:
@bind("sparse_matrix")
class SparseMatrix:
@bind("create_sparse")
def __init__(
self,
nrows: int[Positive],
ncols: int[Positive],
) -> None: ...
@bind("sparse_multiply")
@contract(
pre=lambda c: c.args.x.shape == (c.self.ncols,) and c.args.x.dtype == "float64",
post=lambda c: c.result.shape == (c.self.nrows,) and c.result.dtype == "float64",
)
def multiply(
self,
x: Float64Vector[From(np.ndarray), CPUResident],
) -> Float64Vector: ...This allows:
- semantic API redesign
- Pythonic APIs
- decoupling native APIs from exposed APIs
- explicit validation of user-facing expectations
- reusable runtime checks without changing native source code
The framework allows transforming procedural APIs into clean object-oriented APIs.
Native:
call sparse_multiply(A, x, y)Exposed Python API:
y = A.multiply(x)This is called:
semantic API projection.
The projection records how Python-level self, arguments, and return values map to native parameters.
Projection defines how a Python API maps onto native parameters. It does not decide which Python API to expose. Choosing that shape is a separate step, and it should arrive as a proposal the user reviews, not a decision the generator makes silently.
The exact contract remains the ground truth: the complete native interface, with nothing hidden, reordered, or renamed. A style projection is a second artifact derived from it.
native declarations
-> exact semantic contract complete, mechanical, always generated
-> proposed style projection optional, reviewable, checked in
-> user edits
-> wrapper generation
Most idiom decisions follow recognizable native patterns and need nothing more than a rule table:
- a pointer and its companion length become one array argument;
- an output parameter becomes a returned value;
- a status code with a known success value becomes a raised exception;
- an allocate/free pair becomes a class with a lifetime;
- a callback and its
void *ctxbecome one Python callable; - a
get_x/set_xpair becomes a property.
Rule-based projection is reproducible. The same contract and the same style profile always produce the same proposal, so it can run in CI and its output can be diffed.
Other choices are editorial rather than structural: naming, grouping, docstrings, which arguments deserve defaults, whether a family of procedures is really a class. An AI language model can propose those where a rule table cannot.
Two constraints keep that safe:
-
A proposal is an artifact, not a build step. A model runs when a user asks for a proposal. The result is written to a
.pyifile, reviewed, and committed. Wrapper builds never call a model. - Every proposal is checked against the exact contract. A projection is accepted only when it reduces to the exact native call: same symbols, same argument order and types, nothing invented. That checker is deterministic, so an unverifiable proposal is rejected no matter where it came from.
This is the same trust boundary the rest of the system uses. A suggestion may come from anywhere; only completed, verified policy reaches code generation.
There is no single Pythonic form. A numerical kernel, a handle-based device API, and a configuration library each want a different surface, and two projects can reasonably disagree about the same source.
Styles are therefore named, selectable profiles rather than one hardcoded transformation:
| Style | Shape it proposes |
|---|---|
exact |
The native interface, unchanged. Always available as the baseline. |
idiomatic |
Output parameters as results, status codes as exceptions, conventional names. |
object |
Constructor and destructor families as classes with methods and properties. |
array |
Pointer and length pairs as arrays carrying shape and layout contracts. |
A style profile is data: enabled rules, naming conventions, and defaults. Users extend the built-in profiles or define their own, and a project can pin the profile it generated from so later proposals stay stable.
A command-line spelling such as --style idiomatic, with --pythonic selecting the default idiomatic profile, would choose the proposal. It never changes what the exact contract records.
Semantic types represent:
what an object means conceptually.
NOT:
- its memory layout
- its native language representation
- its ABI representation
Examples:
Float64MatrixSparseMatrixTensor3DCSRMatrixComplexVectorDeviceBuffer
Semantic types are language-independent.
Coercions define:
how one type can be adapted into another.
Examples:
int -> floatnp.ndarray -> Float64MatrixTorchTensor -> Float64MatrixCuPyArray -> DeviceBuffer
Coercions should be explicit in the semantic interface so that the runtime can reject surprising conversions and explain accepted ones.
The semantic interface can declare allowed coercions.
Example:
def scale(alpha: float[From(int)], x: Float64Vector) -> Float64Vector: ...Meaning:
int -> float
is an allowed coercion for alpha.
def solve(
A: Float64Matrix[
From(np.ndarray),
ORDER_F,
Writable,
"N", "N",
],
b: Float64Vector[
From(np.ndarray),
"N",
],
) -> Float64Vector["N"]: ...This means:
- target semantic type for
A:Float64Matrix
- allowed coercion for
A:np.ndarray -> Float64Matrix
- required constraints for
A:- Fortran-contiguous (
ORDER_F) - writable
- square shape
- Fortran-contiguous (
- cross-argument contract:
A.shape[0] == A.shape[1] == b.shape[0]
Constraints define:
requirements the final adapted representation must satisfy.
Constraints are NOT coercions.
Examples:
PositiveWritableORDER_FORDER_CCPUResident- shape subscriptions such as
Float64["N", "N"] Aligned(64)FiniteNonNullBounded
A constraint is usually local to one value: dtype, shape, device, alignment, mutability, ownership, or value range.
The runtime coercion engine is responsible for converting accepted Python objects into semantic runtime objects.
Responsibilities:
- find an allowed conversion path from the observed input type to the target semantic type
- rank competing conversion paths by cost, safety, and zero-copy potential
- apply conversions in order
- preserve ownership and lifetime metadata
- emit a trace that can be shown in diagnostics
Example conversion trace:
argument A:
np.ndarray(shape=(10, 10), dtype=float64, order=C)
-> copy_to_fortran_order
Float64Matrix(shape=(10, 10), dtype=float64, order=F, owner=temporary)
The runtime validation engine checks that semantic runtime objects satisfy the declared constraints and contract predicates.
Responsibilities:
- validate per-argument constraints after coercion
- validate cross-argument preconditions before native calls
- validate return-value postconditions after native calls
- validate object invariants after mutating methods
- produce structured errors with the failing parameter, expected condition, observed value, and coercion trace
Example validation error:
ValidationError in solve(A, b)
parameter: A
failed: ORDER_F
observed: order='C', shape=(10, 10), dtype=float64
hint: declare From(np.ndarray, copy=True) or pass np.asfortranarray(A)
The validation engine is runtime-oriented. It does not replace static typing; it protects the native ABI boundary and provides clear diagnostics for dynamic Python inputs.
Runtime validation contracts are reusable groups of validation rules that describe the semantic obligations of an API.
Contracts may include:
- preconditions: requirements before a native call
- postconditions: guarantees after a native call
- invariants: requirements that must remain true for an object over its lifetime
- aliasing rules: whether inputs may overlap in memory
- mutation rules: which arguments may be modified
- ownership rules: whether returned objects borrow, own, or view native memory
Example:
@contract(
pre=[
lambda ctx: ctx.args.A.shape == (ctx.args.N, ctx.args.N),
lambda ctx: ctx.args.b.shape == (ctx.args.N,),
lambda ctx: ctx.args.A.device == ctx.args.b.device == 'cpu',
],
post=[
lambda ctx: ctx.result.shape == (ctx.args.N,),
lambda ctx: ctx.result.dtype == "float64",
],
invariants=[
lambda ctx: not ctx.result.aliases(ctx.args.A)",
],
)
def solve(
A: Float64Matrix[From(np.ndarray), CPUResident],
b: Float64Vector[From(np.ndarray), CPUResident],
) -> Float64Vector: ...Contracts are higher-level than constraints. A constraint can say b has shape N; a contract can say A and b agree on the same N and that the returned vector does not alias mutable input storage.
The architecture separates:
| Concept | Meaning |
|---|---|
| Semantic type | what the object is |
| Coercion | how another type becomes it |
| Constraint | local requirements on an adapted value |
| Validation contract | API-level preconditions, postconditions, invariants, and aliasing rules |
| Backend adapter | semantic object → ABI representation |
This separation is fundamental.
Allowed coercions declared in .pyi are implemented through a runtime coercion registry in the equivalent .py file.
Example:
@coercion(np.ndarray, Float64Matrix, implicit=True, cost=1, zero_copy="if_compatible")
def ndarray_to_matrix(A: np.ndarray) -> Float64MatrixObject:
return Float64MatrixObject.from_numpy(A)This registers:
np.ndarray -> Float64Matrix
inside the runtime registry.
Validation contracts can also be registered and reused by name.
Example:
@validation_contract
def square_linear_system(ctx):
A = ctx.arg("A")
b = ctx.arg("b")
result = ctx.result
ctx.require(A.ndim == 2, "A must be a matrix")
ctx.require(A.shape[0] == A.shape[1], "A must be square")
ctx.require(b.shape == (A.shape[0],), "b must match A rows")
ctx.ensure(result.shape == b.shape, "solution shape must match b")The interface can then reference the contract:
@contract(square_linear_system)
def solve(A: Float64Matrix, b: Float64Vector) -> Float64Vector: ...This allows common validation logic to be shared across Fortran, C, C++, Rust, and CUDA backends.
Suppose:
x = solve(np.ones((10, 10)), np.ones(10))Runtime pipeline:
Input objects
↓
Find semantic target types
↓
Find coercion paths
↓
Apply coercions
↓
Validate argument constraints
↓
Validate contract preconditions
↓
Backend adapters
↓
Native ABI call
↓
Validate contract postconditions and invariants
↓
Return Python object
The runtime should support composed coercions.
Example:
TorchTensor
↓
np.ndarray
↓
Float64Matrix
The runtime can automatically infer:
TorchTensor -> Float64Matrix
through graph traversal when the path is declared safe and allowed for the target API.
Coercions may contain metadata.
Example:
@coercion(
np.ndarray,
Float64Matrix,
implicit=True,
cost=1,
zero_copy=True,
preserves_aliasing=True,
)
def ndarray_to_matrix(A):
...Possible metadata:
- implicit/explicit
- safe/unsafe
- cost
- zero-copy
- ownership
- device awareness
- aliasing behavior
- mutability preservation
The runtime should internally use semantic runtime objects.
Example:
class Float64MatrixObject:
ptr: int
shape: tuple[int, int]
strides: tuple[int, int]
owner: object | None
device: str
writable: bool
aliases: set[int]These objects are:
- language-independent
- runtime-oriented
- semantic representations
NOT:
- NumPy arrays
- Fortran descriptors
- Eigen matrices
Backend adapters convert:
Semantic runtime object
↓
Native ABI representation
Examples:
- Fortran descriptors
- Eigen maps
- C structs
- CUDA tensors
Adapters should receive values only after coercion and validation have completed. This keeps ABI code focused on call mechanics instead of user-input cleanup.
The framework should support wrapping:
.so.dll- static libraries
without source code.
Users provide:
- semantic
.pyi - coercions if needed
- validation contracts if needed
- optional metadata
No source parsing required.
Parsers are helpers.
NOT the foundation.
Possible parsers:
- Fortran parser
- C parser
- C++ parser
- Rust parser
Their role:
- generate starter
.pyi - synchronize declarations
- help users bootstrap wrappers
The semantic interface remains canonical.
The framework should support libraries implemented in multiple languages simultaneously.
Example:
- Fortran numerical kernels
- C runtime layer
- C++ object systems
- Rust runtime safety
- CUDA kernels
All unified through:
- semantic types
- coercions
- constraints
- validation contracts
- backend adapters
Suppose:
subroutine solve_system(A, b, x)class Mesh {
public:
void refine();
};extern "C" fn optimize(ptr: *mut f64, len: usize) -> i32;The semantic API may expose:
class Solver:
@contract(pre=square_linear_system)
def solve(
self,
A: Float64Matrix[From(np.ndarray), ORDER_F],
b: Float64Vector[From(np.ndarray)],
) -> Float64Vector: ...
class Mesh:
@contract(post=[lambda ctx:ctx.self.is_valid()])
def refine(self) -> None: ...
def optimize(
x: Float64Vector[From(np.ndarray), Writable, CPUResident],
) -> OptimizationResult: ...The user does not care about implementation language. The semantic layer records type meaning, conversion policy, validation policy, and backend dispatch.
The runtime must manage:
- borrowed references
- owned references
- zero-copy views
- temporary coercions
- destruction policies
- aliasing constraints
- mutation contracts
This is one of the hardest parts of the system.
The runtime should avoid unnecessary copies whenever possible.
Examples:
| Conversion | Strategy |
|---|---|
| NumPy F-order → Fortran | zero-copy |
| NumPy C-order → Fortran descriptor requiring F-order | copy or reject, depending on contract |
| NumPy → Eigen::Map | zero-copy when dtype, alignment, and strides match |
| Torch CUDA → CPU array | copy, unless API accepts GPU memory |
| CuPy array → CUDA kernel | zero-copy when stream and device contracts match |
The runtime should optimize coercion paths automatically while still honoring explicit API contracts.
The architecture is especially useful for:
- HPC
- FEM
- CFD
- climate models
- tensor runtimes
- numerical libraries
- GPU computing
- scientific Python ecosystems
because these domains already contain:
- mixed-language systems
- difficult interoperability
- legacy Fortran/C++ code
- array-heavy APIs
- strict shape, device, ownership, and aliasing requirements
The project should generate custom CPython extensions directly.
Reasons:
- full control over runtime
- full control over arrays
- full control over coercions
- full control over validation contracts
- better diagnostics
- better ownership handling
- better performance
The project should NOT fundamentally depend on:
- SWIG
- ctypes
- pybind11
although optional backends may exist later.
Diagnostics are extremely important.
The framework should provide:
- clear coercion errors
- constraint validation errors
- contract validation errors
- coercion trace visualization
- ownership diagnostics
- backend dispatch diagnostics
Example diagnostic:
ContractError in Solver.solve(A, b)
contract: square_linear_system
failed: b.shape == (A.shape[0],)
observed:
A.shape = (10, 10)
b.shape = (8,)
coercion trace:
A: np.ndarray -> Float64Matrix [zero-copy]
b: np.ndarray -> Float64Vector [zero-copy]
This should be much better than typical SWIG/f2py errors.
Third-party ecosystems should be able to register:
- semantic types
- coercions
- constraints
- validation contracts
- backend adapters
This allows:
- NumPy support
- Torch support
- JAX support
- CUDA support
- sparse matrix ecosystems
- domain-specific runtimes
- Define the
.pyi-style semantic grammar. - Represent semantic types, argument mappings, ownership rules, constraints, and validation contracts in the IR.
- Generate a minimal Python-facing wrapper skeleton from the IR.
- Implement the coercion registry.
- Support direct coercions, composed coercion paths, cost ranking, and zero-copy metadata.
- Add structured coercion traces for diagnostics.
- Implement local constraint validation for shape, dtype, contiguity, device, mutability, ownership, and alignment.
- Attach validation failures to source parameters and semantic declarations.
- Run validation after coercion and before backend adaptation.
- Add reusable contract declarations for preconditions, postconditions, invariants, aliasing, mutation, and ownership.
- Support named contract registration and inline contracts in the semantic interface.
- Validate cross-argument relationships such as matching dimensions, shared devices, non-overlapping buffers, and stable object invariants.
- Include contract traces in diagnostics.
- Implement initial Fortran and C adapters.
- Add C++ and Rust adapters after the semantic runtime is stable.
- Add CUDA/device-memory adapters once device contracts are available.
- Add optional parser frontends that generate starter semantic interfaces.
- Add plugin APIs for NumPy, Torch, JAX, CUDA, and sparse matrix ecosystems.
- Keep parser output editable and subordinate to the canonical semantic interface.
- Define style profiles as data and implement the deterministic rule set.
- Implement the reducibility checker that proves a projection maps back to the exact contract.
- Add assisted proposals as a separate authoring command whose output is a reviewed artifact.
- Keep the exact contract generated and authoritative under every style.
The final system becomes:
- a semantic wrapper compiler
- a runtime interoperability framework
- a mixed-language scientific runtime layer
- a semantic coercion engine
- a runtime validation engine
- a runtime validation contract system
- a modern replacement for old wrapper systems
The key innovation is:
Semantic interoperability
instead of
parser-centric wrapper generation
The architecture is built around:
Semantic API
↓
Coercions
↓
Constraints
↓
Validation contracts
↓
Semantic runtime objects
↓
Backend adapters
↓
Native execution
The project focuses on:
- clean semantic APIs
- runtime interoperability
- mixed-language support
- runtime coercion
- runtime validation
- runtime validation contracts
- scientific computing
- extensibility
- high performance
- language independence
while avoiding:
- parser dependence
- compiler dependence
- rigid ABI-centric designs
- old wrapper system limitations.