QMixer¶

class torchrl.modules.QMixer(state_shape: Union[Tuple[int, ...], Size], mixing_embed_dim: int, n_agents: int, device: Union[device, str, int])[source]¶

QMix 混合器。

通过一个单调的超网络将智能体的局部 Q 值混合成一个全局 Q 值，其参数从全局状态获取。引自论文 https://arxiv.org/abs/1803.11485 。

它将每个智能体所选动作的局部值（形状为 (*B, self.n_agents, 1)）转换为一个全局值（形状为 (*B, 1)）。与 torchrl.objectives.QMixerLoss 一起使用。参阅 examples/multiagent/qmix_vdn.py 获取示例。

参数：

state_shape (tuple or torch.Size) – 状态的形状（不包含潜在的批处理维度）。
mixing_embed_dim (int) – 混合嵌入维度的大小。
n_agents (int) – 智能体的数量。
device (str or torch.Device) – 用于网络的 PyTorch 设备。

示例

>>> import torch
>>> from tensordict import TensorDict
>>> from tensordict.nn import TensorDictModule
>>> from torchrl.modules.models.multiagent import QMixer
>>> n_agents = 4
>>> qmix = TensorDictModule(
...     module=QMixer(
...         state_shape=(64, 64, 3),
...         mixing_embed_dim=32,
...         n_agents=n_agents,
...         device="cpu",
...     ),
...     in_keys=[("agents", "chosen_action_value"), "state"],
...     out_keys=["chosen_action_value"],
... )
>>> td = TensorDict({"agents": TensorDict({"chosen_action_value": torch.zeros(32, n_agents, 1)}, [32, n_agents]), "state": torch.zeros(32, 64, 64, 3)}, [32])
>>> td
TensorDict(
    fields={
        agents: TensorDict(
            fields={
                chosen_action_value: Tensor(shape=torch.Size([32, 4, 1]), device=cpu, dtype=torch.float32, is_shared=False)},
            batch_size=torch.Size([32, 4]),
            device=None,
            is_shared=False),
        state: Tensor(shape=torch.Size([32, 64, 64, 3]), device=cpu, dtype=torch.float32, is_shared=False)},
    batch_size=torch.Size([32]),
    device=None,
    is_shared=False)
>>> vdn(td)
TensorDict(
    fields={
        agents: TensorDict(
            fields={
                chosen_action_value: Tensor(shape=torch.Size([32, 4, 1]), device=cpu, dtype=torch.float32, is_shared=False)},
            batch_size=torch.Size([32, 4]),
            device=None,
            is_shared=False),
        chosen_action_value: Tensor(shape=torch.Size([32, 1]), device=cpu, dtype=torch.float32, is_shared=False),
        state: Tensor(shape=torch.Size([32, 64, 64, 3]), device=cpu, dtype=torch.float32, is_shared=False)},
    batch_size=torch.Size([32]),
    device=None,
    is_shared=False)

mix(chosen_action_value: Tensor, state: Tensor)[source]¶

混合器的前向传播。

参数：: chosen_action_value – Tensor，形状为 [*B, n_agents]
返回值：: Tensor，形状为 [*B]
返回值类型：: chosen_action_value

QMixer¶

文档

教程

资源