Overview
A collection of common utility functions and classes used across Habitat Baselines for observation batching, video generation, image processing, action space handling, and policy distribution helpers.
Description
This module provides a broad set of shared utilities for the Habitat Baselines training infrastructure. Key components include:
- CustomFixedCategorical and CategoricalNet: Categorical distribution wrapper and neural network head for discrete action policies.
- CustomNormal and GaussianNet: Normal distribution wrapper and neural network head for continuous action policies, with configurable standard deviation handling (log std, softplus, clamping, learnable parameter).
- batch_obs: High-performance observation batching that transposes a list of observation dicts into a dict of batched tensors, with CUDA pinned-memory optimization and numpy acceleration.
- generate_video: Utility for generating episode videos and optionally logging them to TensorBoard or saving to disk.
- Image processing functions:
center_crop, image_resize_shortest_edge, tensor_to_depth_images, tensor_to_bgr_images, and get_image_height_width.
- Action space helpers:
get_action_space_info, get_num_actions, is_continuous_action_space, and iterate_action_space_recursively.
- LagrangeInequalityCoefficient: A learnable Lagrange multiplier for constrained optimization in RL training.
- Checkpoint helpers:
get_checkpoint_id and poll_checkpoint_folder.
- Scheduling functions:
linear_decay and cosine_decay for learning rate or parameter scheduling.
Usage
Use these utilities when building or extending Habitat Baselines training pipelines. The observation batching (batch_obs) is critical for efficient rollout collection. The distribution classes (CategoricalNet, GaussianNet) serve as action heads in actor-critic policies. Video generation and image utilities are used during evaluation. The Lagrange coefficient is used for constrained policy optimization objectives.
Code Reference
Source Location
Signature
class CategoricalNet(nn.Module):
def __init__(self, num_inputs: int, num_outputs: int) -> None: ...
def forward(self, x: Tensor) -> CustomFixedCategorical: ...
class GaussianNet(nn.Module):
def __init__(self, num_inputs: int, num_outputs: int, config: "DictConfig") -> None: ...
def forward(self, x: Tensor) -> CustomNormal: ...
def batch_obs(
observations: List[DictTree],
device: Optional[torch.device] = None,
) -> TensorDict: ...
def generate_video(
video_option: List[str],
video_dir: Optional[str],
images: List[np.ndarray],
episode_id: Union[int, str],
checkpoint_idx: int,
metrics: Dict[str, float],
tb_writer: TensorboardWriter,
fps: int = 10,
verbose: bool = True,
keys_to_include_in_name: Optional[List[str]] = None,
) -> str: ...
class LagrangeInequalityCoefficient(nn.Module):
def __init__(
self,
threshold: float,
init_alpha: float = 1.0,
alpha_min: float = 1e-4,
alpha_max: float = 1.0,
greater_than: bool = False,
): ...
def lagrangian_loss(self, x): ...
Import
from habitat_baselines.utils.common import (
batch_obs,
generate_video,
CategoricalNet,
GaussianNet,
LagrangeInequalityCoefficient,
linear_decay,
cosine_decay,
center_crop,
image_resize_shortest_edge,
get_action_space_info,
get_num_actions,
)
I/O Contract
Inputs (batch_obs)
| Name |
Type |
Required |
Description
|
| observations |
List[DictTree] |
Yes |
List of observation dictionaries from multiple environments
|
| device |
Optional[torch.device] |
No |
Target torch device for the batched tensors; if None, tensors remain on original device
|
Outputs (batch_obs)
| Name |
Type |
Description
|
| return |
TensorDict |
Dictionary mapping sensor names to batched tensors on the target device
|
Inputs (generate_video)
| Name |
Type |
Required |
Description
|
| video_option |
List[str] |
Yes |
List containing "disk" and/or "tensorboard" to specify output targets
|
| video_dir |
Optional[str] |
No |
Directory path to save video file when "disk" is in video_option
|
| images |
List[np.ndarray] |
Yes |
List of RGB image frames to compose the video
|
| episode_id |
Union[int, str] |
Yes |
Episode identifier used in the video filename
|
| checkpoint_idx |
int |
Yes |
Checkpoint index used in the video filename
|
| metrics |
Dict[str, float] |
Yes |
Performance metrics included in the video filename
|
| tb_writer |
TensorboardWriter |
Yes |
TensorBoard writer for video upload
|
| fps |
int |
No |
Frames per second for the generated video (default 10)
|
| verbose |
bool |
No |
Whether to print verbose output (default True)
|
| keys_to_include_in_name |
Optional[List[str]] |
No |
Metric keys to include in filename; if None, all metrics are included
|
Outputs (generate_video)
| Name |
Type |
Description
|
| return |
str |
The generated video filename, or empty string if no images provided
|
Usage Examples
Batching Observations
import torch
from habitat_baselines.utils.common import batch_obs
# observations is a list of dicts, one per environment
observations = [{"rgb": torch.randn(256, 256, 3), "depth": torch.randn(256, 256, 1)} for _ in range(4)]
# Batch observations and move to GPU
batched = batch_obs(observations, device=torch.device("cuda:0"))
# batched["rgb"].shape -> torch.Size([4, 256, 256, 3])
Using CategoricalNet for Discrete Actions
from habitat_baselines.utils.common import CategoricalNet
# Create a categorical action head
action_head = CategoricalNet(num_inputs=512, num_outputs=4)
features = torch.randn(8, 512) # batch of 8
distribution = action_head(features)
actions = distribution.sample()
log_probs = distribution.log_probs(actions)
Generating Evaluation Video
from habitat_baselines.utils.common import generate_video
video_name = generate_video(
video_option=["disk"],
video_dir="/path/to/videos",
images=rendered_frames,
episode_id=42,
checkpoint_idx=100,
metrics={"spl": 0.85, "success": 1.0},
tb_writer=writer,
fps=30,
)
Related Pages