issue/407 - Release the GIL; preallocated workspace#408
Open
pengcheng888 wants to merge 6 commits into
Open
Conversation
Closed
48 tasks
1c6f631 to
c99ea2a
Compare
Collaborator
Author
|
依赖InfiniCore的pr , InfiniTensor/InfiniCore#1226 |
pengcheng888
commented
Jun 9, 2026
|
|
||
| auto graph = std::get<0>(result->second.compiled); | ||
| auto shared_output = std::shared_ptr<InfinilmModel::Output>(new InfinilmModel::Output{std::get<1>(result->second.compiled)->logits->resume_from_blob_()}); | ||
| // Reuse the GraphTensor output captured at compile time. |
Collaborator
Author
There was a problem hiding this comment.
nvidia会存在double free 的问题。定位到是graph也给Tensor一个deleter, 导致了二次释放。 修改了shared_output后,不再有double free 的问题。
Collaborator
There was a problem hiding this comment.
为什么会二次释放?如果你删掉,会不会导致有些地址无法释放?
pengcheng888
commented
Jun 10, 2026
| * Slots may overlap; scratch_bytes is max span, not sum of slots. Safe use requires | ||
| * temporal reuse across forward phases. | ||
| */ | ||
| class WorkspaceManager { |
Collaborator
Author
There was a problem hiding this comment.
WorkspaceManager的名字来源于vllm
新增了WorkspaceManager,去管理forward的buffer。
(1) 在MLP等模块构造函数中会调用register_buffer函数,提供buffer信息。 (不会分配gpu)
(2) 模型创建结束后,调用finalize_and_bind函数,根据汇总的total_bytes 申请一个DataType::U8的大的scratch_buffer_。(此时才分配空间)
(3) 推理过程中会复用这个scratch_buffer_,不再申请tensor
PanZezhong1725
requested changes
Jun 10, 2026
| size_t num_kv_heads, | ||
| size_t layer_idx); | ||
| size_t layer_idx, | ||
| const infinicore::Device &device); |
Collaborator
There was a problem hiding this comment.
device为什么所有的构造器都要传
|
|
||
| auto graph = std::get<0>(result->second.compiled); | ||
| auto shared_output = std::shared_ptr<InfinilmModel::Output>(new InfinilmModel::Output{std::get<1>(result->second.compiled)->logits->resume_from_blob_()}); | ||
| // Reuse the GraphTensor output captured at compile time. |
Collaborator
There was a problem hiding this comment.
为什么会二次释放?如果你删掉,会不会导致有些地址无法释放?
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Motivation
Closes #
Type of Change
feat— new feature / new modelfix— bug fixperf— performance improvement (no behavioral change)refactor— code restructuring without behavior changetest— adding or fixing tests onlydocs— documentation onlybuild/ci— build system or CI configurationchore— tooling, formatting, or other non-code changesTest Results of Involved Models on Supported Platforms (Please attach screenshots)
Benchmark / Performance Impact
Notes for Reviewers
Checklist
Title, Branch, and Commits
feat(nvidia): …,fix(cuda/gemm): …).<type>/xxx-yyyy-zzzzwhere<type>matches the PR title's Conventional Commits type and words are joined with hyphens (seeCONTRIBUTING.md§Branches).CONTRIBUTING.md§Pull Requests).main— the branch is rebased cleanly on top of the currentmain.fixup!/squash!/wipcommits remain.Scope and Design
CONTRIBUTING.md§Code/General).printf/std::cout/print(...)left behind, orTODOwithout an owner and issue link.General Code Hygiene (applies to all languages)
CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General).the `seqlens_k` tensor) (CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General; §Python).C++ Specific (if C++ files changed)
CONTRIBUTING.md§C++).CONTRIBUTING.md§C++).new/delete; RAII / smart pointers / existing allocators are used.scripts/format.py.csrc/models/llama_legacy/.Python Specific (if Python files changed)
CONTRIBUTING.md§Python).CONTRIBUTING.md§Python).scripts/format.py.python/infinilm/auto_config.py.Testing
examples/test_infer.py), or specify the reason for skipping.examples/bench.py), or specify the reason for skipping.test/bench/test_benchmark.py), or specify the reason for skipping.python/infinilm/server/inference_server.py+scripts/test_perf.py), or specify the reason for skipping.Build, CI, and Tooling
Documentation
README.md,CONTRIBUTING.md, or inline docs updated when behavior, build flags, or developer workflow changed.!orBREAKING CHANGE:footer.Security and Safety