If you are experiencing problems building or using FTorch please see below for guidance on common problems or queries.
The reason input and output tensors to/from
torch_model_forward are contained in arrays
is because it is possible to pass multiple input tensors to the forward()
method of a torch net, and it is possible for the net to return multiple output
tensors.
The nature of Fortran means that it is not possible to set an arbitrary number of inputs to the torch_model_forward subroutine, so instead we use a single array of input tensors which can have an arbitrary length. Similarly, a single array of output tensors is used.
Note that this does not refer to batching data. This should be done in the same way as in Torch; by extending the dimensionality of the input tensors.
torch.inference_mode(), torch.no_grad(), or torch.eval() somewhere like in PyTorch?By default we disable gradient calculations for tensors and models and place models in
evaluation mode for efficiency.
These can be adjusted using the requires_grad and is_training optional arguments
in the Fortran interface. See the API procedures documentation
for torch_tensor_from_array and
torch_model_load etc. for details.
FTorch makes heavy use of Fortran interfaces to module procedures to achieve
overloading of subroutines
such that for users do not need to call a different subroutine for each rank or type
of tensor.
If you make a call to a subroutine that fails to match anything in the interface you will face a compile-time error of the form:
42 | call torch_tensor_from_array(tensor, in_data, torch_kCPU, permute_dims=tensor_permute)
| 1
Error: There is no specific subroutine for the generic 'torch_tensor_from_array' at (1)
The first thing to do in this instance is to inspect the interface you are trying to call, and instead attempt to call the specific procedure (rank and dtype) you expect to use. This can often provide more instructive error messages about what you are doing incorrectly.
Such errors can also occur if you pass a temporary array where the procedure expects
to receive a Fortran array with the target property. For example:
34 | call torch_tensor_from_array(a, [1.0_wp], torch_kCPU, requires_grad=.true.)
| 1
Error: There is no specific subroutine for the generic ‘torch_tensor_from_array’ at (1)
That is, the second argument should be a Fortran array with the target
property, not the temporary array [1.0_wp]. This kind of thing was possible in
FTorch at v1.0 but has since been removed because it is erroneous. Similarly
for expressions involving torch_tensors, e.g., products such as
34 | call torch_tensor_from_array(a, 1.0*in_data1, torch_kCPU, requires_grad=.true.)
| 1
Error: There is no specific subroutine for the generic ‘torch_tensor_from_array’ at (1)
and slices such as
36 | call torch_tensor_from_array(a, in_data1(1,:), torch_kCPU, requires_grad=.true.)
| 1
Error: There is no specific subroutine for the generic 'torch_tensor_from_array' at (1)
Note
If you recently upgraded FTorch and your call used layout as the third argument,
see the Deprecated signature
section below.
Another possible cause of compile-time errors is passing integers of the wrong
kind. FTorch uses explicit integer kinds from the iso_fortran_env intrinsic module:
tensor_shape and tensor_strides
arguments of the tensor constructors, and the return values of
torch_tensor_get_shape and
torch_tensor_get_stride) use 64-bit
integers (int64).ndims), device_index, and permute_dims
use 32-bit integers (int32). As this is the default integer kind on all
supported platforms, plain integer literals and default integer variables can
be used for these arguments.For example, passing a default kind integer array as tensor_shape will raise:
5 | call torch_tensor_ones(a, 2, [10, 20], torch_kFloat32, torch_kCPU)
| 1
Error: Type mismatch in argument 'tensor_shape' at (1); passed INTEGER(4) to INTEGER(8)
To fix this, declare shape arrays using int64:
use, intrinsic :: iso_fortran_env, only: int64
integer(int64), parameter :: tensor_shape(2) = [10, 20]
Conversely, permute_dims expects a default (32-bit) integer array, so passing
an int64 array will raise the 'no specific subroutine' error described above:
8 | call torch_tensor_from_array(b, in_data, torch_kCPU, permute_dims=perm)
| 1
Error: There is no specific subroutine for the generic 'torch_tensor_from_array' at (1)
Use a default integer array for permute_dims, e.g., permute_dims=[2, 1].
Note
If you recently upgraded FTorch and your call used layout as the third argument,
see the Deprecated signature
section below.
torch_tensor_from_array signatureThe update of FTorch to v2.0 brought in a breaking API change to torch_tensor_from_array. If you see this error after upgrading FTorch:
Error: There is no specific subroutine for the generic 'torch_tensor_from_array' at (1)
and you are passing layout as the third argument after data, the interface
has changed.
The layout argument has been deprecated and superseded by permute_dims,
which appears after device_type, and is optional. The old layout signature and
functionality has been moved to a separate interface
torch_tensor_from_array_legacy, which
is deprecated and will be removed in a future version.
Note
If the error is for a different reason (e.g., passing a temporary array, expression, or slice as an argument, or mismatching integer kinds), see the sections above.
The recommended fix is to update your calls to the new signature in one of the following ways:
1) If your Fortran and Torch arrays used the default layout indexing ([1, 2, ..., n])
you can omit layout and permute_dims entirely:
```fortran
! Old
call torch_tensor_from_array(tensor, data, [1, ..., n], torch_kCPU)
! New
call torch_tensor_from_array(tensor, data, torch_kCPU)
2) If you were doing a transpose of a square tensor update your call to use `permute_dims`:fortran
! Old
call torch_tensor_from_array(tensor, data, [n, ..., 1], torch_kCPU)
! New
call torch_tensor_from_array(tensor, data, torch_kCPU, permute_dims=[n, ..., 1])
3) If you rely on unintended behaviour of `layout` then the old call is still available
through [[ftorch_tensor(module):torch_tensor_from_array_legacy(interface)]], but
note that this will be removed completely in the future:fortran
! Old
call torch_tensor_from_array(tensor, data, tensor_layout, torch_kCPU)
! New call torch_tensor_from_array_legacy(tensor, data, tensor_layout, torch_kCPU) ``` 4) If you need deeper control over exactly how you want the data to appear in Torch (the shape and strides) and know what you are doing with memory and array layouts, you can use torch_tensor_from_blob.
Whenever you execute code involving
torch_tensors on each side of an equals sign,
the overloaded assignment operator should be triggered. As such, if you aren't
using the bare use ftorch import then you should ensure you specify
use ftorch, only: assignment(=) (as well as any other module members you
require). See the tensor documentation for more details.
If FTorch unit tests segfault when built with flang, a common cause is an
older pFUnit version which then exposes a bug in flang.
The solution is to use pFUnit v4.18.2 or newer, then reconfigure and rebuild FTorch (and its unit
tests), e.g., specifying to cmake an explicit path to the newer pFUnit (e.g.
-DCMAKE_INSTALL_PREFIX="path/pfunit-4.18.2).
If you are building FTorch with gfortran and are specifying the Fortran 2008
standard (e.g., with the compiler flag -std=f2008 or by default) then you may
get compiler warnings of the form:
Warning: The structure constructor at (1) has been finalized. This feature was removed by f08/0011. Use -std=f2018 or -std=gnu to eliminate the finalization.
These warn that the structure finalizer of the
torch_tensor derived type is triggered when a tensor
goes out of scope, despite the fact that this feature was removed from the 2008
standard. That is, the torch_tensor_delete
subroutine is called so that the associated memory is automatically freed.
Firstly, this is the behaviour that we want so we should not be too concerned.
Secondly, structure finalizers are not used anywhere in FTorch, so we believe
this warning to be errorneous. Use of the structure constructor for the
torch_tensor type would be something like
program
use, intrinsic :: iso_c_binding, only: c_null_ptr
use ftorch
implicit none
type(torch_tensor) :: tensor
tensor = torch_tensor(c_null_ptr)
end program
While this code would compile successfully, the warning mentioned above would be raised.
Warning
The code snippet above is not the intended way to create a tensor. The intended way is to use the provided API procedures such as torch_tensor_from_array or torch_tensor_ones. The code snippet above is only intended to illustrate the use of the structure constructor and the associated warning.
See the tensor documentation for more
details on the memory management of tensors and the use of the finalizer. For
technical details on f08/0011, we refer to
https://wg5-fortran.org/N2001-N2050/N2006.txt.