Skip to content

feat: rebuild #112-114 on current main (AWS placement defaults, manifest-creation fix, k8s deploy_provision dispatch) - #116

Merged
brettchien merged 6 commits into
mainfrom
feat/k8s-onboarding-rebuild-112-114
Aug 28, 2026
Merged

feat: rebuild #112-114 on current main (AWS placement defaults, manifest-creation fix, k8s deploy_provision dispatch)#116
brettchien merged 6 commits into
mainfrom
feat/k8s-onboarding-rebuild-112-114

Conversation

@brettchien

Copy link
Copy Markdown
Contributor

Summary

Rebuilds #112, #113, #114 as one clean PR against current main — same reason as #115: this stack's branches were cut before #105-110's content landed, so merging #112 in place would hit the same squash-merge/dirty-base cascade #106 hit. Rebuilding proactively this time instead of discovering it mid-merge.

The 6 original commits (3 from #112, 1 from #113, 2 from #114) cherry-picked with two small conflicts, both resolved (detail below) — no semantic changes, purely mechanical reconciliation with #115's already-merged content.

Conflicts resolved (both trivial, both expected)

  1. crates/studio-cp/src/lib.rs test module — feat(studio-cp,oabctl): build a default manifest when none exists yet — the actual fix for studio#111 #113's original commit and feat(studio-cp,oab-mcp): list_aws_profiles / list_k8s_contexts tools (studio#104) #105/feat: rebuild #106-110 on current main (list_namespaces/list_service_accounts, expected_principal, console provider picker+wiring, fleetsK8sToml.ts) #115's content both appended a test at the same insertion point (end of mod tests). Resolved by keeping both sets of tests (nothing removed).
  2. crates/studio-cp/src/lib.rs K8sFleetBinding.expected_principalfeat(studio-cp,oabctl): k8s deploy_provision dispatch (studio#104, resumed after #111) #114's branch (like feat(oabctl): expose create.rs's AWS-placement defaults for reuse (studio#111) #112/feat(studio-cp,oabctl): build a default manifest when none exists yet — the actual fix for studio#111 #113) re-added this field because it was cut before feat(studio-cp,oab-mcp): K8sFleetBinding.expected_principal + k8s_fleet_config_write (studio#104) #107 merged (feat(studio-cp,oab-mcp): K8sFleetBinding.expected_principal + k8s_fleet_config_write (studio#104) #107's version landed via feat: rebuild #106-110 on current main (list_namespaces/list_service_accounts, expected_principal, console provider picker+wiring, fleetsK8sToml.ts) #115). Identical field, two different doc-comment wordings. Kept feat(studio-cp,oab-mcp): K8sFleetBinding.expected_principal + k8s_fleet_config_write (studio#104) #107's (canonical, already-merged) version. This is the exact conflict flagged in feat(studio-cp,oabctl): k8s deploy_provision dispatch (studio#104, resumed after #111) #114's own commit message and in K8s fleet onboarding — console UX for provider selection ('+ New fleet') #104's comment thread — expected, not a surprise.

Verification

Rust side: same aws-sdk-ec2 OOM constraint as every PR in this effort — hand-verified the merge result (no duplicate struct/fn definitions, tools()'s Tool::new( count still matches the catalog_advertises_the_named_tools test's assert_eq!(names.len(), 16), t_provision's provider-dispatch block present). CI is the gate.

With this, deploy_provision fully supports k8s end-to-end (#104's provisioning-side scope) and AWS's "+ New fleet" can create a genuinely new agent, not just redeploy one created via oabctl create (#111's scope).

Ref #104, #111.

…udio#111)

First sub-item of #111 ("+ New fleet" first-creation gap). oabctl create's
interactive wizard collects VPC/subnet/security-group among other things
before it can build a manifest and apply it — but the VPC/subnet/SG parts
are already effectively "sensible defaults", not real interactive choices:
select_subnets already auto-picks (private+NAT > private > public, up to 3
AZs) with zero prompting, and "Create new (oab-{name})" is the SG wizard's
own default suggestion.

Makes `list_vpcs`/`VpcInfo`, `select_subnets`/`SubnetInfo`,
`list_security_groups`/`SgInfo` pub (was module-private, module itself was
`mod create` — now `pub mod create`), and adds one new function,
`default_security_group`, extracting the "create new (oab-{name}), reuse if
it already exists" logic `run()` has inline into something callable without
the interactive prompt path.

No behavior change to `run()`/`oabctl create` — pure visibility change plus
one new function built from existing, already-working logic (the AWS calls
in `default_security_group` are the same calls `run()`'s inline SG-create
branch already makes).

This is groundwork only — nothing calls these yet. Next: the console-facing
provision path needs a VPC choice too (unlike subnet/SG, there's no
zero-prompt default for "which VPC" today — `run()` always asks). Then wiring
these into a "build a default manifest, apply via deploy_apply's path"
function is the actual fix for #111.

Local build/test hit the known aws-sdk-ec2 OOM constraint on this machine
(same pre-existing environment issue documented in prior PRs' descriptions) —
change is a mechanical visibility change + one function built from
already-proven inline logic, hand-verified against the diff. CI is the gate.

Ref: studio#111.
… has exactly one default (studio#111)

Follow-up on this same PR: VpcInfo now carries is_default (was folded into
the human-readable label only, not usable programmatically). default_vpc()
picks the account/region's default VPC when there's exactly one — the
closest zero-prompt equivalent to what select_subnets/default_security_group
already give for subnet/SG, since run()'s wizard never had a non-interactive
default for VPC choice at all.

Deliberately errors (not a heuristic guess) when there's zero or more than
one default VPC — a caller with an explicit VPC choice (e.g. future per-fleet
config, mirroring how fleets.toml already carries region/profile) should
skip this and pass that VPC straight to select_subnets/default_security_group
instead.

Still groundwork — nothing calls this yet.

Ref: studio#111.
… VPC/subnet/SG defaults (studio#111)

Follow-up on this same PR. Wraps default_vpc + select_subnets +
default_security_group behind one function that takes an aws_config::
SdkConfig, not an Ec2Client — so studio-cp (which doesn't depend on
aws-sdk-ec2 directly) can reach it without adding that dependency, keeping
Ec2Client an oabctl-internal detail (same "RuntimeDriver is the only layer
with vendor terms" boundary ADR-2 already established elsewhere).

Still groundwork — nothing calls this yet. Next: build_default_manifest in
studio-cp, calling this + spec defaults (resources 256/512, empty secrets,
FARGATE/X86_64), then wire provision_from_library to branch create-vs-
redeploy based on whether load_manifest finds a stored manifest.

Ref: studio#111.
… — the actual fix for #111

provision_from_library (backing deploy_provision / the console's "+ New
fleet" flow) now checks load_manifest first: if a manifest is already
stored, behavior is unchanged (redeploy patches image/bundle_from and
re-applies). If none exists, it builds a fresh OABServiceManifest
(build_default_manifest) using #112's zero-prompt defaults (VPC/subnet/SG
via default_networking, resources 256/512, FARGATE/X86_64, empty secrets —
see #111's comment thread for why empty secrets is valid and not a gap this
needs to solve) and applies it via provision_manifest, a new oabctl::
studio_api helper that mirrors provision() but takes a structured manifest
instead of YAML text (keeps serde_yaml an oabctl-internal detail).

This is the actual fix: the console's "+ New fleet" wizard can now create a
genuinely new agent end-to-end, not just redeploy an agent someone already
created via the CLI. Once merged, k8s's deploy_provision dispatch (#104)
lands on this same branch point — load_manifest is provider-agnostic (S3
key, not ECS-specific), so the create-vs-redeploy check doesn't need to
change for k8s; only the "build a fresh manifest" + "apply it" halves need
a Runtime::Kubernetes(...) branch alongside this Runtime::Ecs(...) one.

configFrom for the new manifest points at artifacts/{ns}/{name}/config.toml
— the same key Bundle::artifact_objects already uploads a copy of the
composed config.toml to, and the same convention oabctl create's wizard
uses. Unit-tested (default_config_from_uri_matches_artifact_objects_key).

Manually verified every field against crates/oabctl/src/manifest.rs's
struct definitions (OABServiceManifest/Metadata/Spec/Resources/Runtime/
EcsRuntime/EcsNetworking) and OABServiceManifest::validate()'s requirements
(apiVersion "oab.dev/v2", kind "OABService", CPU "256" is in
VALID_ECS_CPU, capacityProvider "FARGATE" is valid) — this is the riskiest
change this session (real infra creation), so more care than usual went
into checking it by hand given the environment's known aws-sdk-ec2 OOM
constraint prevented a local cargo check. CI is the gate.

Ref: studio#111.
…sumed after #111)

Resumes #104's k8s deploy_provision work now that #111 (#112/#113) gave both
drivers a shared, provider-agnostic create-vs-redeploy branch point —
load_manifest is just an S3 key lookup, it doesn't care which Runtime
variant a stored manifest holds.

- oabctl::studio_api::provision_k8s: provision_manifest's k8s counterpart,
  applies through K8sDriver instead of EcsDriver. Takes two separate
  credential contexts (aws_config for the S3 bundle carrier — hooks.pre_seed
  is provider-agnostic, still S3 regardless of runtime — and a kubeconfig
  context for the actual apply) since k8s provisioning genuinely needs both
  simultaneously, unlike the ECS path where one SdkConfig covers everything.
- studio-cp::build_default_k8s_manifest: Runtime::Kubernetes counterpart to
  build_default_manifest. No VPC/subnet/SG (ECS-only networking concept);
  service_account comes from K8sFleetBinding.expected_principal when it
  names one (k8s_service_account_from_principal extracts the bare name from
  the system:serviceaccount:<ns>:<name> form the New Fleet wizard's service-
  account picker writes — KubernetesRuntime.service_account wants the bare
  name, verified against k8s_driver.rs's own
  build_deployment_wires_service_account_and_node_selector test).
- studio-cp::provision_from_library_k8s: provision_from_library's k8s
  counterpart. Compose/bundle-upload logic is duplicated rather than shared
  for now (deliberate — avoids reworking provision_from_library's shape
  again while it's still unmerged; worth revisiting once both paths are
  proven). Redeploy preserves the stored manifest's k8s runtime config,
  only bumps the image, same guarantee the AWS path gives.

Nothing calls provision_from_library_k8s yet — oab-mcp's deploy_provision
tool still needs a provider param to dispatch to it (the OabMcp struct also
has no k8s-fleet-binding awareness yet to resolve context/expected_principal
from). That wiring is the next piece.

Manually verified every field against manifest.rs's struct definitions,
same care as #113 given the environment can't locally compile (aws-sdk-ec2
OOM) — CI is the gate.

Ref: studio#104, studio#111.
…104)

t_provision now branches on a new provider arg (default "aws", unchanged
behavior): "k8s" calls provision_from_library_k8s (#114) with context/
expected_principal taken as direct args, not resolved via a fleet name —
the console's identity form already collects context/namespace/service-
account directly (#108/#109), and for a brand-new fleet there's no existing
K8sFleetBinding to look up anyway. Fleet-scoped k8s lookup (for redeploying
into an already-known k8s fleet) is left as explicit future work, not
needed for this dispatch to exist.

Also updates deploy_provision's tool description, which was stale after
provider/context/expected_principal params documented.

NOTE — branch lineage: this stack (#112#113#114→this) was cut from `main`
before #104's original stack (#105-110) merged, not from #110 — so
K8sFleetBinding.expected_principal (originally #107) is re-added here too.
Identical field in both places; trivial merge conflict to resolve whenever
both land, flagging explicitly rather than silently duplicating without a
note.

With this, deploy_provision fully supports k8s end-to-end (provisioning
side) — the remaining piece for full onboarding is unblocking console's
k8s "Next" button (#108/#109's placeholder), a separate follow-up.

Ref: studio#104.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant