Skip to content
This repository was archived by the owner on Mar 23, 2023. It is now read-only.

connection failure #207

@lhj-git

Description

@lhj-git

🐛 Describe the bug

I found a runtime error while running the code:
The client socket has failed to connect to any network address of (hcp-bb-03, 52873). The client socket has failed to connect to hcp-bb-03:52873 (errno: 110 - Connection timed out)
using command line :colossalai run --nproc_per_node 4 --master_port 29505 train.py

Environment

image

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type
    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions