Skip to content

Cannot get more than 8% GPU utilization on M4 with 48 gb (even using --resident-budget 30 GiB for example) #2

Description

@ericlsimplifi

tensorfold serve "$MODEL"
--served-name qwen233-a35b
--resident-budget 39GiB
--loader-backend native
--host 127.0.0.1
--port 8421

GPU stays 1- -max- 8 % (most of the time closer to 1 %)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions