roboto.experimental.ingest#
Describe what an uploaded file carries, so the platform can register it without opening the file.
A caller declares the topics a file contributes data to, what each of them carries, and which files a read
of each opens; the types that compose those files into sessions live in roboto.experimental.sessions.
Submodules#
Package Contents#
- roboto.experimental.ingest.DeclaredTimelineSource#
One timeline source a file’s topic data carries, with the bounds it spans in this file.
- class roboto.experimental.ingest.Field(/, **data)#
Bases:
pydantic.BaseModelOne column of a topic’s data, identified by name, type, and unit.
pathlists the names from the schema root down to this field, so a nested field’s path extends its parent’s.nameandpathstate one fact twice (the last path element is the field’s name), so either may be omitted and derives from the other: a top-level field needs onlyname, and a nested field needs onlypath. A vector column (e.g. a LeRobotobservation.statefeature) is expressed as a parent field holding the array plus one child field per named element.This is what a caller declares.
SchemaFieldRecordis what the platform returns for a field it has stored, and carries the identifiers it assigns.- Parameters:
data (Any)
- canonical_data_type: roboto.domain.topics.record.CanonicalDataType#
Roboto’s normalized type for the field, used for cross-format reads and visualization.
- data_type: str#
Native type of the field as recorded by the source format (e.g.
"float32").
- model_config#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- name: str = ''#
Name of the field. Defaults to the last element of
pathwhen onlypathis given; at least one ofnameandpathmust be declared.
- path: list[str]#
The names from the schema root down to this field; a nested field’s path extends its parent’s path. Omitted,
None, or empty defaults to[name]; when given, every element must be non-empty and the last element must equalname.
- class roboto.experimental.ingest.FileTopicDeclaration(/, **data)#
Bases:
TopicDeclarationOne topic a file contributes data to, declared on the file itself rather than inside a session.
Adds to
TopicDeclarationthe wall-clock instant the data was captured at, which a topic declared inside a session takes from the entry enclosing it.- Parameters:
data (Any)
- anchor_ns: roboto.time._EpochNanosecondsFromTime | None#
the real-world time, in nanoseconds since the Unix epoch, at which that data’s time 0 occurred. Must fall after the Unix epoch, and be small enough to fit in the signed 64-bit integer the platform stores it in. Also accepts any
roboto.time.Timeat runtime, read asroboto.time.to_epoch_nanoseconds()reads it; convert with that function first to satisfy a type checker. When omitted, the data keeps the anchor it already carries from an earlier declaration, and keeps an offset of 0 when it carries none: its timestamps read exactly as declared.An anchor covers the whole slice named by
data_range, not this topic alone, so every topic declared over that slice moves to it; two declarations sharing a slice may not name different anchors.- Type:
Optional wall-clock anchor for the data this declaration names
- roboto.experimental.ingest.MAX_FILES_AND_TOPICS_PER_REQUEST = 500#
Cap on how many files and topics one request may name, counting each file entry and each topic declaration in it.
Each file and topic named costs the platform another round of writes, and a request has a fixed amount of time to finish them all, so a request naming more than this is refused outright rather than left to run out of time; split a larger batch across several calls. The request models below and in
roboto.experimental.sessionsapply the cap when the request body is constructed, and the platform checks it again on every call that declares files or topics, so a body built without these models is held to the same cap.
- roboto.experimental.ingest.MESSAGE_ENVELOPE_TIMELINE_SOURCES: dict[str, tuple[roboto.domain.topics.record.TimelineSourceKind, str]]#
Stored kind and stored name of each timeline source a container stamps on its records, keyed by declared
kind.These sources sit in the message envelope rather than in the topic’s data columns. The stored kind says which of the envelope’s two timestamps a source is, log time or publish time, and reads select a source by its stored name.
MCAP log time and MP4 presentation time each name the timeline their container stamps on every record it holds: for MCAP, the instant the recorder wrote the record; for MP4, the instant the frame is shown. The platform stores one such timeline per topic schema, so the two share a stored kind, differ only in the name a read selects them by, and cannot both be declared on one topic.
The request models in this module check each declaration against the kind and name in this table, and the platform stores the declaration under exactly those.
- class roboto.experimental.ingest.McapLogTimeSource(/, **data)#
Bases:
_DeclaredSourceBaseMCAP’s log time: when the recorder wrote each record to the file.
MCAP carries these timestamps in its message envelope rather than in a data column, so this source names no field.
- Parameters:
data (Any)
- kind: Literal['mcap_log_time'] = 'mcap_log_time'#
Discriminator identifying this entry as MCAP log time.
- class roboto.experimental.ingest.McapPublishTimeSource(/, **data)#
Bases:
_DeclaredSourceBaseMCAP’s publish time: when each message was published on the bus.
MCAP carries these timestamps in its message envelope rather than in a data column, so this source names no field.
- Parameters:
data (Any)
- kind: Literal['mcap_publish_time'] = 'mcap_publish_time'#
Discriminator identifying this entry as MCAP publish time.
- class roboto.experimental.ingest.Mp4PresentationTimeSource(/, **data)#
Bases:
_DeclaredSourceBasePresentation time: when each frame is shown, relative to the start of the media.
These timestamps sit in the message envelope rather than in a data column, so this source names no field. The platform stores it as a message-envelope timeline source named
presentation_time, which is the name a read selects it by.Data declared with this source is readable from an MCAP of encoded frames whose messages’ log times are the presentation times, when the topic lists that MCAP in
TopicDeclaration.representations. The MP4 itself cannot be listed, since no storage format a representation can state describes it. A topic with no representations is registered, and a read of it returns no rows, asTopicDeclaration.representationsdescribes.- Parameters:
data (Any)
- kind: Literal['mp4_presentation_time'] = 'mp4_presentation_time'#
Discriminator identifying this entry as MP4 presentation time.
- class roboto.experimental.ingest.RepresentationDeclaration(/, **data)#
Bases:
pydantic.BaseModelOne representation of a topic’s data: a file a read of the topic can open, and how that file holds the data.
Listed in
TopicDeclaration.representationswhen a topic is declared, and inTopicRepresentations.representationswhen a topic’s representations are replaced. The platform does not open the file when it is listed, so what this states is taken on trust, and a read of the topic that picks this representation opensfile_idand decodes it as stated.Every file named by a representation that is untransformed, or whose every transformation is an
encode, must hold the topic’s rows:At the same positions in each of those files, so a
TopicDeclaration.data_rangenames the same rows whichever one a read opens. In an MCAP, positions count only the topic’s own messages.Decoding to the topic’s declared schema, carrying the timestamps its timeline sources describe: not rebased, not converted to other units, not rounded.
A representation whose file breaks either is accepted, and reads of it return the wrong rows or fail.
A representation covers the whole topic unless it states a
field_path, in which case it covers that field and everything under it. Its file still holds every row, each with its timestamp: a read that takes fields from several files pairs those files row by row, and fails when their row numbers or timestamps differ.Four things identify a representation among those of one topic: what it covers (the whole topic, or one field), its
storage_format, itscontent_formatand itstransformations. Over each part of a file it is declared on, a topic holds at most one representation per combination of the four: two with the same combination cannot be listed together, and one declared later takes the earlier one’s place.A read decodes a
PARQUETrepresentation’s file as Parquet; the file needs no.parquetextension. A read can decode anMCAPrepresentation’s file when all three of the following hold:The file is chunked and carries a summary section indexing those chunks.
Exactly one of the file’s channels carries the topic’s name.
That channel’s schema is in an encoding the platform can decode:
ros1msg,ros2msg,ros2idl,omgidl,jsonschema, orjson.
A recording holding several topics is therefore read one topic at a time, each from the channel carrying its name; the file’s other channels are neither decoded nor checked.
- Parameters:
data (Any)
- content_format: str | None = None#
Format of the data inside the container, such as
"jpeg"for re-encoded images or"compressedVideo"for passed-through video frames, orNonewhen unspecified. A read names it throughcontent_format.
- field_path: list[str] | None#
Field of the topic’s schema this representation covers, as the names from the schema root down to it.
None, the same as leaving it out, makes this a representation of the whole topic; an empty list and an empty name are refused.The path must equal the
pathof a field the topic’s schema declares. That field may have fields under it, and the representation then covers those too. A level of the schema that only the paths of deeper fields pass through, with no field declared at it, cannot be named.A read takes each field from the representation covering the deepest field that contains it, so a representation of one field supplies that field whatever its transformations, and the representations of the whole topic supply the rest. A read asking for untransformed data (
RepresentationSelector.raw()) leaves out every representation that states a transformation.
- file_id: str#
the file the topic is declared on, when its own bytes hold the data, or another file of the org, such as a per-topic MCAP converted out of a PX4 ULog. Its status must be
Available, and the caller must be able to edit it. A read through the SDK or the web app downloads this file with the reader’s own download permission on it, and Roboto’s AI tools read it for anyone who can read the topic.- Type:
File a read opens to get the topic’s data
- model_config#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- storage_format: roboto.domain.topics.record.RepresentationStorageFormat#
Container
file_idholds the data in. Within one request, every representation naming a file must state the same format for it, whichever of the request’s topics, files or sessions lists it; a request stating two formats for one file is refused when it is built.
- transformations: list[str]#
The transformations applied to the topic’s data to produce what
file_idholds, in the order applied, as"<kind>:<param>"descriptors such as["downsample:0.7", "encode:jpeg"];TransformationKindlists the kinds. Empty for untransformed data. A read names them throughtransformations, and with no selector the platform prefers the representation with the fewest.A representation can be read by row position when it is untransformed or every transformation is an
encode. One with any other transformation, such as adownsample, cannot, and a read of a topic declared over adata_rangenever uses it: on such a topic, everything it covers must also be covered by representations that can.Within one topic, every representation naming a file must state the same transformations for it: a read that takes two representations from one file and finds their transformations differ raises
RobotoReadPlanExecutionExceptionwith kindinconsistent-scan-tasks-on-file.
- class roboto.experimental.ingest.Schema(/, **data)#
Bases:
pydantic.BaseModelThe structure of one topic’s data: the columns it carries.
However a schema is produced, whether hand-written field by field or converted from a source format’s own metadata, the registered result is the same: schemas are content-addressed server-side. Identity covers every attribute of every field (name, path, source data type, canonical type, and unit), so identical declarations collapse to a single stored schema no matter how many times they are repeated, while declarations differing in any field attribute are stored separately.
A column a timeline source reads is declared by typing it
Timestampwith aTimeUnitunit; nothing else marks it. Which of a topic’s timeline sources reads fall back to is not part of the schema: it is stated ontimeline_sourcesand can be changed later, so the same columns are one schema no matter which source is preferred.This is what a caller declares, so it carries no checksum: the platform computes that from the fields.
TopicSchemaRecordis the stored schema the platform returns, carrying that checksum and the identifiers it assigns.- Parameters:
data (Any)
- fields: list[Field]#
Declared columns of the topic’s data. At least one is required, and every field’s path must be unique within the schema.
- model_config#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- name: str | None = None#
Informational label for the schema (often the topic name). Not part of schema identity.
- class roboto.experimental.ingest.SchemaFieldSource(/, **data)#
Bases:
_DeclaredSourceBaseTimestamps read from a column of the topic’s own data, such as a
timestampfield.- Parameters:
data (Any)
- field_path: list[str]#
the names from the schema root down to that field. The field must be declared on the topic’s schema and typed
Timestamp.- Type:
Path of the field holding this source’s timestamps
- kind: Literal['field'] = 'field'#
Discriminator identifying this entry as timestamps read from a data column.
- class roboto.experimental.ingest.TopicDeclaration(/, **data)#
Bases:
pydantic.BaseModelA topic a file contributes data to.
The platform registers the data from the declaration alone, never opening the file, so the declaration states the topic’s name, the structure of its rows, the bounds of every timeline source those rows carry, and which part of the file they occupy.
List a file’s topic declarations on the
SessionFilefor that file; that entry anchors everything it declares. To register a file’s topic data without naming a session, throughdeclare_topics(), build aFileTopicDeclarationinstead; with no enclosing entry to anchor it, that form carries its own anchor.The file a topic is declared on is the one its data belongs to: the data’s slices, bounds and anchors, and the windows sessions hold it over, are all stated against that file. Which files a read opens to get the data is a separate statement,
representations, so data can belong to a file the platform cannot decode (a PX4 ULog, a ROS.bag, a CSV) and be read from files converted out of it.- Parameters:
data (Any)
- data_range: roboto.domain.topics.record.DataRange | None = None#
startis the first covered position, andendis one past the last. Set this when one file packs a topic’s data into slices, such as a single episode inside a LeRobot v3 data file. Omit it when the data covers whatever encloses this declaration: the slice claimed by theSessionFilecarrying it, or the whole file when the topics are declared on the file itself throughdeclare_topics(). A range set inside aSessionFilemust sit within that entry’s owndata_range.Values are in the file’s own units: stored-row positions (counted from 0), or nanoseconds of the file’s media time for video. Each file uses exactly one of the two, so the pair needs no unit marker; which one applies follows from the file’s format. A read applies the range to the file it opens, and it opens only files named by representations that can be read by row position, as
RepresentationDeclaration.transformationsdescribes: each of those files must hold the same rows at the same positions. Over a range, everything a topic’s representations cover must be covered by ones that can be read by row position; the platform refuses the declaration otherwise.- Type:
The part of the file this topic’s data occupies, as
(start, end)
- model_config#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- representations: list[RepresentationDeclaration]#
The files a read of this topic’s data opens, each with how it holds that data. List the file the topic is declared on here, naming its
file_id, when its own bytes are readable; list other files when the data is read from them, such as the per-topic MCAPs converted out of a PX4 ULog. A topic may list several, and a read picks among them byRepresentationSelector.A topic with no representations is still registered: its data counts toward the bounds of every session holding the file, and a read of it returns no rows, or is refused when the read’s
RepresentationSelectorsets a criterion. A video, or a proprietary log with no file converted out of it, is declared that way.Redeclaring the topic over the same part of the file adds the representations listed to the ones it has. A listed representation takes the place of the stored one that covers what it covers (the whole topic, or the same field) under the same storage format, content format and transformations, and of every stored one that names the same file, whatever that one covers; every other stored representation stays. Redeclaring the topic with the new files a conversion wrote therefore replaces each stored representation with the listed one that differs from it only in its
file_id; redeclaring it with a file uploaded again in another storage format replaces the stored representation that names that file; and resending the same declaration leaves the stored representations as they are. To remove a representation, or to replace a topic’s representations outright, useset_representations().A representation of one field is stored against that field of the schema it was declared under. While one is stored, the platform refuses a declaration of the topic over the same part of the file with a changed
topic_schema, unless the declaration lists a representation that takes the stored one’s place. To change the schema, list the representation of that field again in the same declaration, or remove it first withset_representations().Refused when the model is built:
Two representations covering the same thing (both the whole topic, or the same field) under the same storage format, content format and transformations.
Two representations naming one file under different transformations.
A representation whose
field_pathnames no field oftopic_schema.A
PARQUETrepresentation on a topic declaring any timeline source butSchemaFieldSource: the others take their timestamps from the message envelope, which a Parquet file does not have.
- timeline_sources: list[DeclaredTimelineSource]#
Timeline sources this file’s topic data carries, each with its own bounds. Data timestamped several ways (a message’s publish time and the time the recorder wrote it, say) declares one entry per source; data timestamped one way declares a list of one. Reads pick a source by name, and resolve to the one marked
is_default_for_readswhen they do not.
- topic_name: str#
Topic this file (or slice of it) contributes data to. Topic names are unique within an org; files declaring the same name contribute to the same topic.
- topic_schema: roboto.experimental.ingest.schema.Schema#
Structure of the topic’s data. Repeat the same schema on every file that uses it; the platform stores each distinct schema once, so repetition costs nothing extra.
- class roboto.experimental.ingest.TopicRepresentations(/, **data)#
Bases:
pydantic.BaseModelThe complete set of representations one topic’s data on a file is read from.
Handed to
set_representations(), which leaves the topic, over the part of the filedata_rangenames, with exactly these representations.RepresentationDeclarationstates what the file each one names must hold.- Parameters:
data (Any)
- data_range: roboto.domain.topics.record.DataRange | None = None#
The part of the file over which the topic’s representations are replaced, or
Nonefor data declared over the whole file. It must equal a range the topic is already declared over on the file: the topic’sTopicDeclaration.data_range, or, for a topic that stated none inside aSessionFile, that entry’sdata_range.
- model_config#
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- representations: list[RepresentationDeclaration]#
The representations the topic ends up with; its representations not listed are removed, those of one field included. An empty list removes every representation: the topic stays registered and still counts toward the bounds of every session holding the file, but reads of it return no rows. Two representations may not cover the same thing (both the whole topic, or the same field) under the same storage format, content format and transformations, or name one file under different transformations; both are refused when the model is built.
set_representations()states what the platform refuses.
- topic_name: str#
Topic whose representations are replaced. It must already be declared on the file.