Skip to content

[Question]: Is A6000 supported? #23

Description

@yawzhe

Describe the issue

A6000支持加速推理吗?我们暂时没有A100,其他的服务器什么型号的可以呀

Activity

  1. iofu728 commented on Jul 8, 2024

    @iofu728
    Contributor

    Hi yawzhe,

    Thank you for your support in MInference.

    After reviewing the code, it appears that our decoding stage relies on flash-attn. For the prefilling stage, the three ops used are based on a Triton-implemented version of dynamic sparse flash attention, which does not depend on the flash-attn library. We plan to support a version that does not rely on flash-attn in the future, although this version will have slightly higher latency compared to the one using flash-attn.

  2. changed the title [-][Question]: A6000支持加速推理吗?[/-] [+][Question]: Is A6000 supported?[/+] on Jul 8, 2024
  3. iofu728 commented on Jul 15, 2024

    @iofu728
    Contributor

    Please upgrade MInference to version 0.1.4.post3, which does not depend on flash_attn. You can do this by running the following command:

    pip install minference==0.1.4.post3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

feature requestNew feature or requestquestionFurther information is requested

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions