Skip to content

Instances (/instances/)

Instances sit at the core of Shaken Fist's functionality, and are the component which ties most of the other concepts in the API together. Therefore, they are also the most complicated part of Shaken Fist to explain. This description is broken into basic functionality -- showing information about instances -- and then moves onto more advanced topics like creation, deletion, and other lifecycle events.

Note

For a detailed reference on the state machine for instances, see the developer documentation on object states.

Fetching information about an instance

There are two main ways to fetch information about instances -- you can list all instances visible to your authenticated namespace, or you can collect information about a specific instance by providing its UUID or name.

Info

Note that the amount of information visible in an instance response will change over the lifecycle of the instance -- for example when you first request the instance be created versus when the instance has had its disk specification calculated.

REST API calls
Python API client: get all visible instances
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
instances = sf_client.get_instances()
print(json.dumps(instances, indent=4, sort_keys=True))
Python API client: get information about a specific instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
i = sf_client.get_instance('317e9b70-8e26-46af-a1c4-76931c0da5a9')
print(json.dumps(i, indent=4, sort_keys=True))

Instance creation

Instance creation is by far the most complicated call in Shaken Fist in terms of the arguments that it takes. The code in the Python command line client is helpful if you need a fully worked example of every possible permutation. The OpenAPI documentation at https://openapi.shakenfist.com/#/instances/post_instances provides comprehensive and up to date documentation on all the arguments to the creation call.

The instance creation API call also takes three data structures: the diskspec; the networkspec; and the videospec. These structures are not well documented in the OpenAPI interface, so are documented here instead.

Since v0.8 all three are described to the API's input validation layer instead of being handed to the handler unexamined, which has two consequences for every one of them. A key which is not listed below is refused with a 400 naming it, where it used to be silently discarded -- typing siz instead of size got you a default sized disk and no indication that anything had gone wrong. And a value of the wrong type is refused with a 400 naming the key and the element it is in (disk[0].size: Not a valid integer.), where many of them used to reach code which could not cope and produce a server error naming nothing.

Any of the values below may also be sent as JSON null, which means "not supplied" and gets you the same behaviour as omitting the key -- the shipped client and the shipped ansible collection both do that for several of them on an ordinary request. The exceptions are the three values a handler requires: a networkspec's network_uuid, and a videospec's model and memory. For those a null is refused exactly as an omission is, which is the same statement as the rest of this paragraph rather than a different one -- omitting them is already a 400.

diskspec

A diskspec consists of the following fields as a JSON dictionary:

  • size (integer): the size of the disk in gigabytes. Must be zero or greater; a negative size is refused with a 400, because it corrupts the cluster's capacity accounting. Omit this value, or send null, for a disk the size of its base image.
  • base (string): the base image for the disk. This can be a variety of URL-like strings, as documented on the artifacts page in the user guide. For a blank disk, omit this value, send null, or send the literal string "none" -- all three mean the same thing here.
  • bus (enum): the hardware bus the disk device should be attached to on the instance. In general you shouldn't care about this and can omit this value. However, in some cases, such as unmodified Microsoft Windows images it is required. The options available here are: sata; scsi; usb; virtio (the default); and nvme. While ide was previously supported, that support was removed in v0.7 due to extremely poor performance, and a bus of ide is refused with a 400 along with every other value outside that list.
  • type (enum): the type of device, either "disk" (the default) or "cdrom". Any other value is refused with a 400. Before v0.8 any string was accepted here and handed to libvirt as a device name, which meant that libvirt's own device types were accepted and then behaved as plain disks, because only "cdrom" is treated specially by Shaken Fist.

A diskspec must ask for something: one which specifies neither a size nor a base -- including a size of zero with no base, and a base of the literal string "none", which means no base -- is refused with a 400, because it describes a disk nobody asked for.

A full example of a diskspec is therefore:

{
    'size': 20,
    'base': 'debian:11',
    'bus': None,
    'type': None
}

networkspec

Similarly, a networkspec consists of the following fields in a JSON dictionary:

  • network_uuid (string, required): the network the interface should exist on. Despite the name this is not restricted to being a UUID -- the unique name of a network in the namespace you are creating in is accepted as well, which is why this value is published as a plain string rather than as a UUID. A networkspec with no network_uuid, or with a null one, is refused with a 400.
  • macaddress (string): the MAC address of the interface, in the colon separated form 02:00:00:ea:3a:28. Either case is accepted. A value in any other form is rejected with a 400. Omit this value, or send null, to be allocated a MAC address automatically -- which is what the shipped client does on every call.
  • address (string): the IPv4 address to assign to the interface, or the literal string "none" for an interface with no address at all. Omit this value, or send null, to be allocated a random address. Because of the "none" case this value is published as a plain string and is not validated as an IPv4 address.
  • model (string): the model of the network interface card. In general you should not have to set this, although it can matter in some cases, such as unmodified Microsoft Windows images. virtio is the default and is almost always the right answer; e1000, rtl8139, pcnet and the i825xx family are the usual choices for a guest without virtio drivers. There is deliberately no enumeration of legal values here, and a value this API does not recognise is not refused: the set which actually works is whatever the hypervisor's qemu build supports, which varies by node and by release and which this API cannot know. If a model is wrong you will find out when the instance fails to start, not when you create it. The one check the value does get is the published pattern ^[A-Za-z0-9._-]{1,32}$, which every model name fits: the value is rendered into the instance's libvirt domain XML, so a value carrying an XML metacharacter is refused with a 400 (and is escaped at render time regardless).
  • float (boolean): whether to associate a floating IP with this interface to enable external accessibility to the instance. Note that you can float and unfloat an interface after instance creation if desired. A JSON boolean is the expected form, and the shipped client and ansible collection both send one. A range of string spellings (true/false, yes/no, on/off, 1/0, and some single-letter and case variants of each) are also read with the meaning they carry, so "false" does not float the interface. A spelling outside those sets -- "tRue" is one -- is refused with a 400 under the shipped enforce default, but read as true under the warn/off rollback, where the handler's own bool() fallback answers instead of the schema check. Do not rely on either; send a real JSON boolean.

The same structure is passed to POST /instances/{instance_ref}/interfaces when hot plugging an interface into a running instance, where it is a single dictionary rather than a list of them, and is checked identically.

videospec

A videospec differs from a diskspec and a networkspec in that it is not passed as a list. You only have one videospec per instance. Once again, a videospec is a JSON dictionary with the following fields:

  • model (string): the model of the video card to attach to the instance. cirrus is the default, and vga and qxl are the other usual choices -- qxl is the one to pair with SPICE. As with a networkspec's model there is deliberately no enumeration of legal values, and an unrecognised one is not refused, for the same reason: the value is rendered into the libvirt domain XML, so the working set belongs to the hypervisor's qemu build rather than to this API. It carries the same ^[A-Za-z0-9._-]{1,32}$ pattern as the networkspec's model, for the same reason.
  • memory (integer): the amount of video RAM the video card should have, in kibibytes (blocks of 1024 bytes). A value with a fractional part is refused with a 400.
  • vdi (enum): the VDI protocol to use. Options are "vnc", "spice" (the default), "spiceconcurrent", or "spicedebug". spice and spiceconcurrent are the same except that spiceconcurrent allows limited multi-user sessions, with subsequent sessions not experiencing full VDI functionality. spicedebug behaves like spice but also sets G_MESSAGES_DEBUG=all in the qemu environment so the SPICE server emits verbose logs to the qemu log, which is useful for diagnosing connectivity issues. This enumeration is Shaken Fist's own, rather than the hypervisor's, so unlike the two model fields a value outside it is refused with a 400 when you create the instance. It previously reached the instance and then failed in a console request, long after the call which accepted it.

If you supply a videospec at all it must contain both model and memory, or the call is refused with a 400 -- and a null counts as not containing it, since a stored null reaches the hypervisor and produces an instance which cannot start. Omit the whole structure, send null, or send an empty dictionary to get the defaults instead. A null vdi is defaulted like an absent one.

REST API calls
Python API client: create and then delete a simple instance
from shakenfist_client import apiclient
import time

sf_client = apiclient.Client()
i = sf_client.create_instance(
    'example', 1, 1024, None,
    [{
        'size': 20,
        'base': 'debian:11',
        'bus': None,
        'type': 'disk'
    }],
    None, None)

time.sleep(30)

i = sf_client.delete_instance(i['uuid'])

Adding network interfaces after instance creation

As of Shaken Fist v0.8, it is also possible to add network interfaces to an existing instance, assuming that your guest operating system supports device hot plugging (all modern Linux versions do).

REST API calls
Python API client: hot plug a network interface

Note that this example assumes the instance is running an image with the Shaken Fist in guest agent installed.

```python from shakenfist_client import apiclient

sf_client = apiclient.Client()

Create a network to hot plug to

hotnet = sf_client.allocate_network('10.0.0.0/24', True, True, 'hotplug')

...

Hot plug the interface in

netdesc = { 'network_uuid': hotnet['uuid'], 'address': '10.0.0.5', 'macaddress': '02:00:00:ea:3a:28' } sf_client.add_instance_interface(inst['uuid'], netdesc)

Other instance lifecycle operations

A variety of other lifecycle operations are available on instances, including deletion, and power management.

The power management actions available are:

  • soft reboot: gracefully request a reboot from the instance operating system via ACPI. This is not guaranteed to actually work, but if it does is much less likely to cause disk corruption on the instance.
  • hard reboot: the equivalent of holding the reset switch down on a physical machine until it reboots without operating system involvement.
  • power on: turn the instance on, as if the power switch was pressed. Since v0.8 power on operations have the side effect of creating the config drive if one is specified by the instance configuration. That is, you can recreate the config drive by powering the instance off and then on again.
  • power off: turn the instance immediately off, as if the power switch was held down on a physical machine.
  • pause: suspend execution of the instance, but leave it hot in RAM ready to restart.
  • unpause: unsuspend execution of the instance.
REST API calls
Python API client: create and then delete a simple instance
from shakenfist_client import apiclient
import time

sf_client = apiclient.Client()
i = sf_client.create_instance(
    'example', 1, 1024, None,
    [{
        'size': 20,
        'base': 'debian:11',
        'bus': None,
        'type': 'disk'
    }],
    None, None)

time.sleep(30)

i = sf_client.delete_instance(i['uuid'])
Python API client: attempt a soft reboot, and hard reboot if required

Note that this example assumes the instance is running an image with the Shaken Fist in guest agent installed.

import time
from shakenfist_client import apiclient
import sys

sf_client = apiclient.Client()
i = sf_client.create_instance(
    'example', 1, 1024, None,
    [{
        'size': 20,
        'base': 'debian:11',
        'bus': None,
        'type': 'disk'
    }],
    None, None, side_channels=['sf-agent'])

# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
    events = sf_client.get_instance_events(i['uuid'])
    print('Waiting for the instance to start: %s' % events[0]['message'])
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

# Check the instance is created correctly
if i['state'] != 'created':
    print('Instance is not in a created state!')
    sys.exit(1)
print('Instance is created')

# Wait for the agent to report the reboot time
while not i['agent_system_boot_time']:
    print('Waiting for agent to start: %s' % i['agent_state'])
    time.sleep(20)
    i = sf_client.get_instance(i['uuid'])

initial_boot = i['agent_system_boot_time']
print('Instance booted at %d' % initial_boot)

# Now try to soft reboot the instance, wait up to 60 seconds for a reboot to
# be detected
sf_client.reboot_instance(i['uuid'], hard=False)
print('Soft rebooting instance')
time.sleep(60)
i = sf_client.get_instance(i['uuid'])

# Wait for the agent to report the reboot time again
while not i['agent_system_boot_time']:
    print('Waiting for agent to start: %s' % i['agent_state'])
    time.sleep(20)
    i = sf_client.get_instance(i['uuid'])

if i['agent_system_boot_time'] != initial_boot:
    print('Boot time changed from %d to %s'
        % (initial_boot, i['agent_system_boot_time']))

else:
    # We failed to soft reboot, let's hard reboot instead
    sf_client.reboot_instance(i['uuid'], hard=True)
    print('Instance did not reboot, hard rebooting')

Sample output:

$ python3 example.py
Waiting for the instance to start: schedule complete
Instance is created
Waiting for agent to start: not ready (no contact)
Waiting for agent to start: not ready (no contact)
Waiting for agent to start: not ready (no contact)
Instance booted at 1684404969
Soft rebooting instance
Boot time changed from 1684404969 to 1684405036.0
Python API client: power off and then on an instance
import time
from shakenfist_client import apiclient
import sys

sf_client = apiclient.Client()
i = sf_client.create_instance(
        'example', 1, 1024, None,
        [{
            'size': 20,
            'base': 'debian:11',
            'bus': None,
            'type': 'disk'
        }],
        None, None)

# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
    events = sf_client.get_instance_events(i['uuid'])
    print('Waiting for the instance to start: %s' % events[0]['message'])
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

# Check the instance is created correctly
if i['state'] != 'created':
    print('Instance is not in a created state!')
    sys.exit(1)
print('Instance is created')

# Check the instance is created correctly
if i['power_state'] != 'on':
    print('Instance is not in powered on state!')
    sys.exit(1)

# Power the instance off
sf_client.power_off_instance(i['uuid'])
while i['power_state'] != 'off':
    print('Waiting for the instance to power off')
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

time.sleep(30)

# Power the instance on
sf_client.power_on_instance(i['uuid'])
while i['power_state'] != 'on':
    print('Waiting for the instance to power on')
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

print('Done')
Waiting for the instance to start: set attribute
Instance is created
Waiting for the instance to power off
Waiting for the instance to power on
Done
Python API client: pause and unpause an instance
import time
from shakenfist_client import apiclient

sf_client = apiclient.Client()
sf_client.pause_instance('foo')

time.sleep(30)

sf_client.unpause_instance('foo')

Other instance information

We can also request other information for an instance. For example, we can list the instance's network interfaces, or the events for the instance. See the user guide for a general introduction to the Shaken Fist event system.

REST API calls
Python API client: list network interfaces for an instance

Note that the interface details for an instance wont be populated until the instance has started being created on the hypervisor node. Specifically, this can be some time later if an image needs to be fetched from the Internet and transcoded. Therefore in this example we wait for the instance to be created before displaying interface details.

import json
from shakenfist_client import apiclient
import time

sf_client = apiclient.Client()
i = sf_client.create_instance(
        'example', 1, 1024, None,
        [{
            'size': 20,
            'base': 'debian:11',
            'bus': None,
            'type': 'disk'
        }],
        None, None)

# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
    events = sf_client.get_instance_events(i['uuid'])
    print('Waiting for the instance to start: %s' % events[0]['message'])
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

# Check the instance is created correctly
if i['state'] != 'created':
    print('Instance is not in a created state!')
    sys.exit(1)
print('Instance is created')

# Fetch and display interface details
ifaces = sf_client.get_instance_interfaces(i['uuid'])[0]
print(json.dumps(ifaces, indent=4, sort_keys=True))
$ python3 example.py
Waiting for the instance to start: Fetching required blob ffdfce7f-728e-4b76-83c2-304e252f98b1, 30% complete
Instance is created
[
    {
        "floating": null,
        "instance_uuid": "d512e9f5-98d6-4c36-8520-33b6fc6de15f",
        "ipv4": "10.0.0.6",
        "macaddr": "02:00:00:73:18:66",
        "metadata": {},
        "model": "virtio",
        "network_uuid": "6aaaf243-0406-41a1-aa13-5d79a0b8672d",
        "order": 0,
        "state": "created",
        "uuid": "b1981e81-b37a-4176-ba37-b61bc7208012",
        "version": 3
    }
]
Python API client: list events for an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
interfaces = sf_client.get_instance_events('c0d52a77-0f8a-4f19-bec7-0c05efb03cb4')
print(json.dumps(interfaces, indent=4, sort_keys=True))

Note that events are returned in reverse chronological order and are limited to the 100 most recent events.

[
    ...
    {
        "duration": null,
        "extra": {
            "cpu usage": {
                "cpu time ns": 357485828000,
                "system time ns": 66297716000,
                "user time ns": 291188112000
            },
            "disk usage": {
                "vda": {
                    "actual bytes on disk": 956301312,
                    "errors": -1,
                    "read bytes": 406776320,
                    "read requests": 12225,
                    "write bytes": 2105954304,
                    "write requests": 3657
                },
                "vdb": {
                    "actual bytes on disk": 102400,
                    "errors": -1,
                    "read bytes": 279552,
                    "read requests": 74,
                    "write bytes": 0,
                    "write requests": 0
                }
            },
            "network usage": {
                "02:00:00:1d:24:ae": {
                    "read bytes": 147084732,
                    "read drops": 0,
                    "read errors": 0,
                    "read packets": 16484,
                    "write bytes": 2166754,
                    "write drops": 0,
                    "write errors": 0,
                    "write packets": 13144
                }
            }
        },
        "fqdn": "sf-2",
        "message": "usage",
        "timestamp": 1685229509.9592097,
        "type": "usage"
    },
    ...
]

Out-of-band interactions with instances

Shaken Fist supports three types of instance consoles, which provide out-of-band management of instances -- that is, the instance does not need to have functioning networking for these consoles to work. You can read a general introduction of Shaken Fist's console functionality in the user guide. This page focuses on the API calls which are used to implement the console functionality in the Shaken Fist client.

  • Read only console: to download the most recent portion of the read only text serial console, or clear the console, use the /instances/{instance_ref}/consoledata API calls below.
  • Interactive serial console: lookup the console port from the instance details fetch (as described above), and then connect to that port on the hypervisor node with a TCP client such as telnet.
  • Interactive VDI console: lookup the VDI console port from the instance details fetch (as described above), and then connect to that port on the hypervisor with the correct client (currently one of VNC or SPICE). Alternatively, use the /instances/{instance_ref}/vdiconsolehelper API call described below to download a virt-viewer configuration file and then connect with virt-viewer. See the example below for more details.
  • Proxied VDI console: since v0.8, if the cluster operator has enabled the Kerbside integration, use the /instances/{instance_ref}/vdiconsoleproxy API call below to mint a short lived signed token and receive a proxy URL of the form <KERBSIDE_URL>/sf-console.vv?token=<jwt>. The user's viewer redeems this URL against the Kerbside proxy, which validates the token offline and relays the SPICE session, so no direct network access to the hypervisor is needed. See the VDI console tokens operator guide for how the integration is enabled and operated.
REST API calls
Python API client: connect seamlessly to a VDI console using virt-viewer
import os
from shakenfist_client import apiclient
import subprocess
import tempfile
import time

sf_client = apiclient.Client()
i = sf_client.create_instance(
        'example', 1, 1024, None,
        [{
            'size': 20,
            'base': 'debian:11',
            'bus': None,
            'type': 'disk'
        }],
        None, None)

# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
    events = sf_client.get_instance_events(i['uuid'])
    print('Waiting for the instance to start: %s' % events[0]['message'])
    time.sleep(5)
    i = sf_client.get_instance(i['uuid'])

# Check the instance is created correctly
if i['state'] != 'created':
    print('Instance is not in a created state!')
    sys.exit(1)
print('Instance is created')

# We don't use NamedTemporaryFile as a context manager as the .vv file
# will also attempt to clean up the file.
(temp_handle, temp_name) = tempfile.mkstemp()
os.close(temp_handle)
try:
    with open(temp_name, 'w') as f:
        f.write(sf_client.get_vdi_console_helper(i['uuid']))

    p = subprocess.run('remote-viewer %s' % temp_name, shell=True)
    print('Remote viewer process exited with %d return code' % p.returncode)
finally:
    if os.path.exists(temp_name):
        os.unlink(temp_name)

Executing commands within an instance

Since v0.7, assuming a given instance has the Shaken Fist agent installed and running, and was created with a sf-agent side channel, you can use the Shaken Fist agent to move data into and out of the instance and execute commands without the instance needing to have working networking configured. You can read more about the exact requirements for agent connectivity in the API reference guide for agent operations.

Agent Operations are not created directly -- they are a side effect of a call to one of the API methods below, which create an Agent Operation so the caller can track the state of their request. At the time of writing, you can perform the following operations via the agent:

  • copy the contents of a blob into an instance and change its file permissions. The python API client has a helper to upload the file into a blob before copying to the instance.
  • execute a command and return its results (exit code, stdout, stderr).
  • get the contents of a file within an instance into a blob.
REST API calls

Bounding how long an agent operation may take

All three creating calls accept an optional deadline_seconds, and agent/get and agent/put additionally accept an optional progress_timeout_seconds. Both are counts of seconds, and both refuse a negative value with a 400. Both are also capped by the operator ceiling AGENT_OPERATION_MAX_DEADLINE (86400 seconds, one day, unless the operator has changed it), published as the parameters' maximum in the API specification and likewise refused with a 400 above it.

deadline_seconds is how long the operation may continue to be dispatched or execute, counted from the moment the API server received your request rather than from when the agent picked the work up. Time spent queued behind another operation on the same instance, and any preflight work such as fetching a blob onto the hypervisor, both count against it.

progress_timeout_seconds is how long the operation may go without making forward progress. It applies to commands which can report progress -- the transfers behind agent/get and agent/put -- and is the more useful of the two for a large file, because a transfer can be perfectly healthy and still take a long time.

Each parameter has three meanings, and the difference between the first two matters:

You send What happens
nothing The server default applies: AGENT_OPERATION_DEFAULT_DEADLINE (600 seconds) or AGENT_OPERATION_DEFAULT_PROGRESS_TIMEOUT (30 seconds), unless your operator has changed them.
0 None at all. The operation is not bounded by that mechanism.
a positive number That many seconds, up to AGENT_OPERATION_MAX_DEADLINE.

Sending 0 is not the same as omitting the parameter. Streaming a very large file out of an instance is the case the distinction exists for: send deadline_seconds: 0 with a progress_timeout_seconds, and the transfer is allowed to take as long as it takes while a genuine stall is still detected.

There is one exception to 0 meaning unbounded: an operation with both budgets disabled -- deadline_seconds: 0 on agent/execute, which always stores no progress timeout, or combined with an explicit progress_timeout_seconds: 0 on the transfers -- would otherwise hold its instance's single executor slot for as long as the guest command cared to run, blocking every other agent operation against that instance. Such an operation is expired AGENT_OPERATION_MAX_DEADLINE seconds after it last changed state instead.

agent/execute does not accept progress_timeout_seconds, and sending it is a 400. Nothing an executed command does is observable as progress, so a timeout there could never fire; only deadline_seconds is meaningful.

What happens when a budget runs out

An operation which exhausts either budget, and for which no retry is possible (see below), moves to expired, a terminal state distinct from error. error means the operation itself failed; expired means a budget you set ran out, and the operation's expiry_reason field says which: deadline or progress. A deadline expiry is answered by asking for a longer deadline; a progress expiry means the agent stalled, which a bigger deadline does not fix. A wall-clock deadline running out is always final -- there is no time left for a further attempt to deliver anything in -- but a progress-timeout stall may be retried a bounded number of times before landing in expired; see Agent Operations in the operator guide for the retry rules and the node-local reaper that backs them up.

Three places enforce this, so an abandoned operation is retired wherever it happens to be sitting: when it reaches the head of the instance's queue, during preflight (either side of any blob copy), and once per second while the executor is running it. Only the executor can enforce progress_timeout_seconds, since it is the only one of the three watching replies arrive.

The reason is recorded as an audit event against both the operation and its instance, and the instance's copy is the one which survives: an expired operation is eventually hard deleted, like a completed one.

Python API client: execute a command on an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
agentop = sf_client.instance_execute('...uuid...', 'cat /etc/os-release')
print(json.dumps(agentop, indent=4, sort_keys=True))

Which would return something along the lines of:

{
    "attempts": 0,
    "commands": [
        {
            "block-for-result": true,
            "command": "execute",
            "commandline": "cat /etc/os-release"
        }
    ],
    "deadline": 1787428090.5,
    "expiry_reason": null,
    "instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
    "last_progress": null,
    "metadata": {},
    "namespace": "vdi",
    "progress_timeout": 0.0,
    "results": {
        "0": {
            "command-line": "cat /etc/os-release",
            "result": true,
            "return-code": 0,
            "stderr": "",
            "stdout": "PRETTY_NAME=\"Debian GNU/Linux 11 (bullseye)\"..."
        }
    },
    "state": "complete",
    "uuid": "93fb538c-84f5-4ff8-83ba-2be5f5f92954",
    "version": 3
}
Python API client: put a file onto an instance via a blob
import os
from shakenfist_client import apiclient
from shakenfist_client import util

sf_client = apiclient.Client()
if not sf_client.check_capability('blob-search-by-hash'):
    blob = None
else:
    # We can cheat here -- if we already have a blob in the cluster with the
    # checksum of the file we're uploading, we can skip the upload entirely and
    # just reuse that blob.
    blob = util.checksum_with_progress(sf_client, 'README.md')

if not blob:
    artifact = util.upload_artifact_with_progress(
        sf_client, 'upload-to-instance', 'README.md', None)
else:
    print('Recycling existing blob')
    artifact = sf_client.blob_artifact(
        'upload-to-instnace', blob['uuid'], source_url=None)
print('Created artifact %s' % artifact['uuid'])

st = os.stat('README.md')
sf_client.instance_put_blob(
        '...instance_ref...', artifact['blob_uuid'], '/tmp/README.md', st.st_mode)

Which would return something along the lines of:

$ python3 /tmp/demo.py
Calculate checksum: 100%|██████████████████████████| 805/805 [00:00<00:00, 13.0MB/s]
Searching for a pre-existing blob with this hash...
Recycling existing blob
Created artifact 3c0a6a83-e9df-46f0-b9a3-819eb16bea23
Python API client: get a file from an instance via a blob
from shakenfist_client import apiclient
import sys

sf_client = apiclient.Client()
op = sf_client.instance_get('...instance_ref...', '/tmp/README.md')
if '0' not in op.get('results', {}):
    print('Results not available.')
    sys.exit(1)

blob_uuid = op['results']['0'].get('content_blob')
if not blob_uuid:
    print('Results did not include content')
    sys.exit(1)

with open('/tmp/README.md', 'wb') as f:
    for chunk in sf_client.get_blob_data(blob_uuid):
        f.write(chunk)

Fetching information about an Instance's Agent Operations

Additionally, you can list the agent operations for a given instance.

REST API calls
Python API client: get all agent operations for a specific instance

The instance here had the following command line commands run before this sample script was run:

$ sf-client instance upload ...uuid... README.md /tmp/README.md
$ sf-client --simple instance execute ...uuid... "cat /tmp/README.md"
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
agentops = sf_client.get_instance_agentoperations('...uuid...', all=True)
print(json.dumps(agentops, indent=4, sort_keys=True))

Note the all argument here. By default you are only returned agent operations which are queued to execute. To see all agent operations including those which have completed execution, pass all=True. This script outputs:

[
    {
        "attempts": 0,
        "commands": [
            {
                "blob_uuid": "09306f15-b1b3-4850-afb4-f4179559fa7f",
                "command": "put-blob",
                "path": "/tmp/README.md"
            },
            {
                "command": "chmod",
                "mode": 33188,
                "path": "/tmp/README.md"
            }
        ],
        "deadline": 1787428090.5,
        "expiry_reason": null,
        "instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
        "last_progress": null,
        "metadata": {},
        "namespace": "vdi",
        "progress_timeout": 30.0,
        "results": {
            "0": {
                "path": "/tmp/README.md"
            },
            "1": {
                "path": "/tmp/README.md"
            }
        },
        "state": "complete",
        "uuid": "343049d7-da2a-46f2-bb5c-edb783ec1fb9",
        "version": 3
    },
    {
        "attempts": 0,
        "commands": [
            {
                "block-for-result": true,
                "command": "execute",
                "commandline": "cat /tmp/README.md"
            }
        ],
        "deadline": 1787428090.5,
        "expiry_reason": null,
        "instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
        "last_progress": null,
        "metadata": {},
        "namespace": "vdi",
        "progress_timeout": 0.0,
        "results": {
            "0": {
                "command-line": "cat /tmp/README.md",
                "result": true,
                "return-code": 0,
                "stderr": "",
                "stdout": "...content of file..."
            }
        },
        "state": "complete",
        "uuid": "5a00d6f3-19b6-42bc-b1df-ddc4e5a299e9",
        "version": 3
    }
]

Console screen captures

Since v0.8, Shaken Fist has provided an API for collecting screen captures of the instance console. This works for either serial consoles or graphical consoles, its literally the same was whatever would have been displayed on the monitor if the instance was a physical machine.

REST API calls

This API call returns a blob UUID, you then need to collect the contents of the blob using the GET /blobs/{blob_uuid}/data API call. The python Shaken Fist API client perfoms both operations for you and returns an iterator of binary chunks ready for you to process or write to a file.

Python API client: collect a screenshot for an instance an instance
from shakenfist_client import apiclient

sf_client = apiclient.Client()
with open(destination, 'wb') as f:
    for chunk in sf_client.get_screenshot(instance_ref):
        f.write(chunk)

Object References

Instance API responses include references_to and references_from fields that show the relationships between instances and other objects in the system. These fields help you understand how instances are connected to blobs and other objects.

The references_to field shows what objects reference this instance (typically empty for instances). The references_from field shows what blobs this instance references (e.g., disk blobs, NVRAM template blobs).

Example references_from output for an instance
"references_from": {
    "disk": [
        {
            "source_object_type": "instance",
            "source_uuid": "d51aa352-368c-484c-9e4c-4542927b4277",
            "relationship": "disk",
            "relationship_value": "0",
            "target_object_type": "blob",
            "target_uuid": "5117f778-b214-4184-8358-f2c7376b76db",
            "created": 1683995934.357137,
            "last_active": 1684054381.217045
        }
    ],
    "nvram_template": [
        {
            "source_object_type": "instance",
            "source_uuid": "d51aa352-368c-484c-9e4c-4542927b4277",
            "relationship": "nvram_template",
            "relationship_value": null,
            "target_object_type": "blob",
            "target_uuid": "abc123-def456-ghi789",
            "created": 1683995934.357137,
            "last_active": 1684054381.217045
        }
    ]
}

Metadata

All objects exposed by the REST API may have metadata associated with them. This metadata is for storing values that are of interest to the owner of the resources, not Shaken Fist. Shaken Fist does not attempt to interpret these values at all, with the exception of the instance affinity metadata values. The metadata store is in the form of a key value store, and a general introduction is available in the user guide.

Info

Note that for affinity metadata to be processed by the scheduler, it must be present in the instance create API call, which is why that call takes a metadata argument. Adding affinity metadata after instance creation will not affect the placement of that instance, but would affect the placement of future instances.

REST API calls
Python API client: set metadata on an instance
from shakenfist_client import apiclient

sf_client = apiclient.Client()
sf_client.set_instance_metadata_item(instance_uuid, 'foo', 'bar')
Python API client: get metadata for an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
md = sf_client.get_instance_metadata(instance_uuid)
print(json.dumps(md, indent=4, sort_keys=True))
Python API client: delete metadata for an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
sf_client.delete_instance_metadata_item(instance_uuid, 'foo')

Cluster operations

Since v0.8, cluster operations for a given instance have been exposed. This lists the various pieces of queued work that Shaken Fist has executed on a given object, with the limitation that completed cluster operations are hard deleted after CLEANER_DELAY seconds (which defaults to one hour).

REST API calls
Python API client: get queued cluster operations for an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
md = sf_client.get_instance_queued_cluster_operations(img_uuid)
print(json.dumps(md, indent=4, sort_keys=True))
Python API client: get all cluster operations for an instance
import json
from shakenfist_client import apiclient

sf_client = apiclient.Client()
md = sf_client.get_instance_queued_cluster_operations(img_uuid, all=True)
print(json.dumps(md, indent=4, sort_keys=True))

📝 Report an issue with this page