fix(serve): remove IC data-source collapse hack from recommendation deploy - #6101
Open
ZealSV wants to merge 7 commits into
Open
fix(serve): remove IC data-source collapse hack from recommendation deploy#6101ZealSV wants to merge 7 commits into
ZealSV wants to merge 7 commits into
Conversation
Lokiiiiii
suggested changes
Jul 24, 2026
ZealSV
force-pushed
the
remove-ic-sdkt-hack
branch
from
July 27, 2026 18:32
a322ca1 to
e161c32
Compare
…eploy Inference Components now support AdditionalModelDataSources natively (kernel tuning / speculative decoding channels), so _deploy_recommendation no longer needs to collapse an optimized recommendation's base_model + draft channels into a single primary ModelDataSource. Deploy the recommendation's ModelPackage directly and let the hosting stack resolve the channels. Removes the base_model-promotion / OPTION_SPECULATIVE_DRAFT_MODEL rewiring block and its now-unused shape imports (AdditionalModelDataSource, ModelDataSource, S3ModelDataSource, SPECULATIVE_DRAFT_MODEL). Updates the speculative-decoding unit tests to assert the ModelPackage pass-through.
ZealSV
force-pushed
the
remove-ic-sdkt-hack
branch
from
July 27, 2026 20:37
b68d882 to
81bee88
Compare
…onent Adds an end-to-end integration test that builds a model carrying speculative-decoding / kernel-tuning AdditionalModelDataSources (base + draft channels) and deploys it as an Inference Component via ModelBuilder.deploy(inference_config=ResourceRequirements(...)). Asserts the Inference Component reaches InService and that the deployed model still carries the additional sources (guards against a silent client-side collapse). Marked slow_test + gpu_intensive.
The test built the IC model via ModelBuilder(model_path=<s3_uri>), which raises 'Cannot detect required model or inference spec' — a raw S3 path is not a buildable model spec. Build via ModelBuilder.from_jumpstart_config so the container/framework resolve, then attach additional_model_data_sources (carried onto the model through _prepare_container_def_base) before deploying as an Inference Component. Also switch the IC instance to g4dn.xlarge (a GPU is needed for the vLLM/LMI container, not for the 0.6B size) and trim the module docstring.
…down deploy(wait=True) waits for the endpoint, but the Inference Component is created with wait=False, so the IC can still be Creating when deploy() returns. Add _wait_for_ic_terminal to poll the IC to InService/Failed before the status assertion, and wait it out of Creating before teardown (an IC in Creating cannot be deleted, which would strand its GPU endpoint).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Inference Components now support AdditionalModelDataSources natively (kernel tuning / speculative decoding channels), so _deploy_recommendation no longer needs to collapse an optimized recommendation's base_model + draft channels into a single primary ModelDataSource. Deploy the recommendation's ModelPackage directly and let the hosting stack resolve the channels.
Removes the base_model-promotion / OPTION_SPECULATIVE_DRAFT_MODEL rewiring block and its now-unused shape imports (AdditionalModelDataSource, ModelDataSource, S3ModelDataSource, SPECULATIVE_DRAFT_MODEL). Updates the speculative-decoding unit tests to assert the ModelPackage pass-through.
Issue #, if available:
Description of changes:
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.