Martin Kroeker [Sun, 19 Apr 2020 12:55:31 +0000 (14:55 +0200)]
Update getarch.c
Martin Kroeker [Sun, 19 Apr 2020 11:52:58 +0000 (13:52 +0200)]
Update getarch.c
Martin Kroeker [Sun, 19 Apr 2020 11:22:19 +0000 (13:22 +0200)]
Rename the FORCE entries for 24K and 1004K to include the MIPS prefix
Martin Kroeker [Sun, 19 Apr 2020 05:21:48 +0000 (07:21 +0200)]
Disable RPCC macro on MIPS24K
Martin Kroeker [Sun, 19 Apr 2020 04:54:52 +0000 (06:54 +0200)]
Update README.md
Martin Kroeker [Sun, 19 Apr 2020 04:51:57 +0000 (06:51 +0200)]
Update TargetList.txt
Martin Kroeker [Sun, 19 Apr 2020 04:50:51 +0000 (06:50 +0200)]
Add compiler options for MIPS32 24K/1004K
Martin Kroeker [Sat, 18 Apr 2020 21:50:23 +0000 (23:50 +0200)]
rename 1004K, 24K to MIPS1004K, MIPS24K to avoid identifier naming problem
Martin Kroeker [Sat, 18 Apr 2020 19:16:49 +0000 (21:16 +0200)]
Typo fix in MIPS24K addition
Martin Kroeker [Sat, 18 Apr 2020 19:10:18 +0000 (21:10 +0200)]
Add MIPS24K support
Martin Kroeker [Sat, 18 Apr 2020 19:09:32 +0000 (21:09 +0200)]
Handle MIPS24K like P5600
and allow enforcing TARGET=1004K as well (omission from earlier 1004K merge and later introduction of TARGET check)
Martin Kroeker [Sat, 18 Apr 2020 19:07:14 +0000 (21:07 +0200)]
Merge pull request #47 from xianyi/develop
rebase
Martin Kroeker [Thu, 16 Apr 2020 09:45:32 +0000 (11:45 +0200)]
Merge pull request #2563 from zelong-1024/develop
[OpenBLAS]: benchmark error of potrf
l00536773 [Thu, 16 Apr 2020 02:55:10 +0000 (10:55 +0800)]
[OpenBLAS]: benchmark error of potrf
[description]: when the matrix size goes higher than 5800 during the cpotrf test, error info, such as "Potrf info = 5679", will be returned on ARM64 and x86 machines. Uplo = L & F.
[solution]: changed the func for building the matrix so that the complex Hermitian matrix can stay positive definite during the computation.
[dts]:
Martin Kroeker [Wed, 15 Apr 2020 18:23:43 +0000 (20:23 +0200)]
Merge pull request #2557 from martin-frbg/dronebadge
Update and reformat README
Martin Kroeker [Wed, 15 Apr 2020 18:23:17 +0000 (20:23 +0200)]
Merge pull request #2556 from martin-frbg/epicdrone
Add a drone.io multithread test for x86_64
Martin Kroeker [Wed, 15 Apr 2020 17:26:12 +0000 (19:26 +0200)]
Restore USE_OPENMP in the x86 thread test
Martin Kroeker [Wed, 15 Apr 2020 15:38:33 +0000 (17:38 +0200)]
Move all 19.04-based jobs back to ubuntu 18.04
Martin Kroeker [Tue, 14 Apr 2020 17:18:35 +0000 (19:18 +0200)]
try x86_64 test without openmp
Martin Kroeker [Tue, 14 Apr 2020 08:53:28 +0000 (10:53 +0200)]
Add drone.io badge, mention EMAG8180 support, reformat the DYNAMIC_ARCH paragraph
Martin Kroeker [Mon, 13 Apr 2020 20:46:12 +0000 (22:46 +0200)]
Add a multithread test for x86_64
Martin Kroeker [Mon, 13 Apr 2020 19:28:59 +0000 (21:28 +0200)]
Merge pull request #2553 from martin-frbg/issue2444
Add a read memory barrier to the traversal of the buffer slot list
Martin Kroeker [Mon, 13 Apr 2020 16:29:56 +0000 (18:29 +0200)]
Merge pull request #2555 from martin-frbg/issue1137
Handle unaligned data in the SSE2 copy kernel
Martin Kroeker [Mon, 13 Apr 2020 13:56:31 +0000 (15:56 +0200)]
ARMV7 does not support DMB ISHLD, use DMB ISH
Martin Kroeker [Mon, 13 Apr 2020 12:58:52 +0000 (14:58 +0200)]
Convert aligned moves to unaligned
should have no performance impact on reasonably modern cpus and fixes occasional crashes in actual user code.
Martin Kroeker [Mon, 13 Apr 2020 10:34:02 +0000 (12:34 +0200)]
Add a read barrier in the traversing of the buffer list
Needed on systems with weak memory ordering - the inferior, partially working fix from #2544 was already removed in #2551
Martin Kroeker [Mon, 13 Apr 2020 10:24:10 +0000 (12:24 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:22:35 +0000 (12:22 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:18:48 +0000 (12:18 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:17:41 +0000 (12:17 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:16:44 +0000 (12:16 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:14:58 +0000 (12:14 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:14:06 +0000 (12:14 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:12:54 +0000 (12:12 +0200)]
Add (empty) read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:11:58 +0000 (12:11 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:10:37 +0000 (12:10 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:09:24 +0000 (12:09 +0200)]
Add read barrier definition
Martin Kroeker [Mon, 13 Apr 2020 10:06:40 +0000 (12:06 +0200)]
Merge pull request #46 from xianyi/develop
rebase
Martin Kroeker [Sun, 12 Apr 2020 20:34:41 +0000 (22:34 +0200)]
Merge pull request #2551 from martin-frbg/issue2538-2
Increase BUFFER_SIZEs and add a safeguard; supply GEMM_R for POWER8/9
Martin Kroeker [Sun, 12 Apr 2020 17:47:02 +0000 (19:47 +0200)]
Fix parameter overflow
Martin Kroeker [Sun, 12 Apr 2020 17:45:36 +0000 (19:45 +0200)]
Add safeguards for sufficient BUFFER_SIZE
Martin Kroeker [Sun, 12 Apr 2020 17:44:48 +0000 (19:44 +0200)]
Increase default BUFFER_SIZE on ARM, ZARCH and newer x86_64, add GEMM_R for POWER8/9
As shown in #2538, default buffersizes on some platforms were smaller than required in memory.c
and the requirement could never be fulfilled for a calculated GEMM_R on PPC given the fomula used
Martin Kroeker [Sun, 12 Apr 2020 17:39:05 +0000 (19:39 +0200)]
Merge pull request #45 from xianyi/develop
rebase
Martin Kroeker [Fri, 10 Apr 2020 22:35:38 +0000 (00:35 +0200)]
Merge pull request #2547 from sharvil/develop
Add API to set thread affinity on Linux.
Martin Kroeker [Fri, 10 Apr 2020 22:35:07 +0000 (00:35 +0200)]
Merge pull request #2541 from bapt/develop
libname: treat FreeBSD and DragonFly like linux and sunos
Martin Kroeker [Fri, 10 Apr 2020 22:34:04 +0000 (00:34 +0200)]
Merge pull request #2548 from gxw-loongson/develop
Add a GENERIC target for 64bit MIPS
Martin Kroeker [Fri, 10 Apr 2020 21:54:40 +0000 (23:54 +0200)]
Merge pull request #2549 from martin-frbg/fixthreadtest
Match thread count in cpp_thread_test to host capability
Martin Kroeker [Fri, 10 Apr 2020 20:06:44 +0000 (22:06 +0200)]
Match thread count to machine capability
gxw [Thu, 9 Apr 2020 11:25:13 +0000 (19:25 +0800)]
Fix compilation problem on loongson platform
Using "make TARGET=GENERIC" on loongson platform will get the following
error messages:
"make[1]: *** No rule to make target 'sgemm_incopy.o', needed by 'libs'"
Add kernel/mips64/KERNEL.generic to slove the problem.
Sharvil Nanavati [Wed, 8 Apr 2020 19:47:41 +0000 (12:47 -0700)]
Add API to set thread affinity on Linux.
Issue: #2545
Martin Kroeker [Wed, 8 Apr 2020 09:04:51 +0000 (11:04 +0200)]
Add another memory barrier for ARM and a multicore test run on ThunderX to help detect such issues (#2544)
* Add another memory barrier in memory.c to prevent races in memory slot allocation
* Add an all-core test on Drone.io's ThunderX platform and modify dgemm_tester to use all 96 cores
Martin Kroeker [Sat, 4 Apr 2020 20:48:53 +0000 (22:48 +0200)]
Merge pull request #44 from xianyi/develop
Add a Z13 build to the Travis configuration (#2542)
Martin Kroeker [Sat, 4 Apr 2020 20:46:58 +0000 (22:46 +0200)]
Merge pull request #43 from martin-frbg/revert-42-z12ci
Revert 42 z12ci to keep forked develop clean
Martin Kroeker [Sat, 4 Apr 2020 20:45:01 +0000 (22:45 +0200)]
Revert "Add IBM Z to Travis configuration (#42)"
This reverts commit
7972beb3754409db9af3c22dbbe7bd8075c09f6e.
Martin Kroeker [Fri, 3 Apr 2020 14:02:11 +0000 (16:02 +0200)]
Add a Z13 build to the Travis configuration (#2542)
* Add IBM Z to Travis configuration
Martin Kroeker [Fri, 3 Apr 2020 13:59:18 +0000 (15:59 +0200)]
Add IBM Z to Travis configuration (#42)
* Add IBM Z to Travis configuration
Baptiste Daroussin [Fri, 3 Apr 2020 04:20:42 +0000 (06:20 +0200)]
libname: treat FreeBSD and DragonFly like linux and sunos
There is no difference in the way libnames are handle between FreeBSD
and linux or sunos. FreeBSD and DragonFly prefers having sonames as well
Martin Kroeker [Thu, 2 Apr 2020 08:32:19 +0000 (10:32 +0200)]
Merge pull request #41 from xianyi/develop
rebase
Martin Kroeker [Thu, 2 Apr 2020 08:30:37 +0000 (10:30 +0200)]
Make ARMV7 compile with xcode and add a CI job for it (#2537)
* Add an ARMV7 iOS build on Travis
* thread_local appears to be unavailable on ARMV7 iOS
* Add no-thumb option for ARMV7 IOS build to get it to accept DMB ISH
* Make local labels in macros of nrm2_vfpv3.S compatible with the xcode assembler
Martin Kroeker [Wed, 1 Apr 2020 18:00:13 +0000 (20:00 +0200)]
Merge pull request #2536 from martin-frbg/recurs
Add "recursive" option for LAPACK builds with ifort or pgfort as well
Martin Kroeker [Wed, 1 Apr 2020 13:39:16 +0000 (15:39 +0200)]
ifort and pgfort need "recursive" for safe compilation of LAPACK as well
Martin Kroeker [Wed, 1 Apr 2020 13:38:07 +0000 (15:38 +0200)]
ifort and pgfort need "recursive" for compiling LAPACK as well
as shown in Reference-LAPACK issue 401 (their PR 403)
Martin Kroeker [Tue, 31 Mar 2020 18:53:13 +0000 (20:53 +0200)]
Merge pull request #2534 from martin-frbg/issue2496
Fix zero initialization for beta=0 case
Martin Kroeker [Tue, 31 Mar 2020 14:53:56 +0000 (16:53 +0200)]
fix initialization to zero in the NEON SGEMM_BETA kernel as well
Martin Kroeker [Mon, 30 Mar 2020 22:21:02 +0000 (00:21 +0200)]
Fix zero initialization for beta=0 case
use immediate initialization instead of multiplication in case register content is a NaN
Martin Kroeker [Mon, 30 Mar 2020 18:15:59 +0000 (20:15 +0200)]
Merge pull request #2520 from wjc404/develop
Fix avx512 sgemm performance bug when ldc is a multiple of 1024
Martin Kroeker [Mon, 30 Mar 2020 18:15:37 +0000 (20:15 +0200)]
Merge pull request #2533 from martin-frbg/gemmdirect2
Use runtime check for AVX512 capability in DYNAMIC_ARCH builds made on SKX
Martin Kroeker [Thu, 26 Mar 2020 20:25:39 +0000 (21:25 +0100)]
Expose the support_avx512 function provided in dynamic.c
Martin Kroeker [Thu, 26 Mar 2020 20:12:56 +0000 (21:12 +0100)]
Use runtime check for AVX512 (sgemm_direct) capability when using DYNAMIC_ARCH
Martin Kroeker [Thu, 26 Mar 2020 20:06:51 +0000 (21:06 +0100)]
Merge pull request #39 from xianyi/develop
rebase
Martin Kroeker [Tue, 24 Mar 2020 14:44:46 +0000 (15:44 +0100)]
Merge pull request #2530 from martin-frbg/dynmsg
Add message highlighting minimum target choice at end of DYNAMIC_ARCH…
Martin Kroeker [Tue, 24 Mar 2020 14:44:27 +0000 (15:44 +0100)]
Merge pull request #2529 from shengyang-3390/dev1
add ctest for drotm and modified ctest for drot.
Martin Kroeker [Mon, 23 Mar 2020 18:35:51 +0000 (19:35 +0100)]
Add message highlighting minimum target choice at end of DYNAMIC_ARCH builds
related to #2526
Martin Kroeker [Mon, 23 Mar 2020 11:47:19 +0000 (12:47 +0100)]
Merge pull request #2527 from martin-frbg/gemmdirect
Avoid calling DIRECT codepath in DYNAMIC_ARCH on non-SKX
shengyang [Sat, 21 Mar 2020 07:58:21 +0000 (15:58 +0800)]
add ctest for drotm and modified ctest for drot.
make sure that test cases cover all code path when kernel uses looping unrolling.
Martin Kroeker [Sun, 22 Mar 2020 13:33:16 +0000 (14:33 +0100)]
Avoid calling DIRECT codepath in DYNAMIC_ARCH on non-SKX
Martin Kroeker [Sun, 22 Mar 2020 00:03:42 +0000 (01:03 +0100)]
Merge pull request #2521 from martin-frbg/cm-avx512
Use proper extension on the avx512 testcase filename
Martin Kroeker [Sat, 21 Mar 2020 17:47:48 +0000 (18:47 +0100)]
Merge pull request #2525 from andreas-schwab/develop
Fix ARCHCONFIG for Neoverse-N1
Andreas Schwab [Sat, 21 Mar 2020 16:33:33 +0000 (17:33 +0100)]
Fix ARCHCONFIG for Neoverse-N1
../config_kernel.h:24:9: warning: missing whitespace after the macro name
24 | #define ARMV8-march armv8.2-a
| ^~~~~
Martin Kroeker [Fri, 20 Mar 2020 22:05:53 +0000 (23:05 +0100)]
Use proper extension on the avx512 testcase filename
The need to call it .tmp existed only when it was generated by a tmpfile call, and the "-x c" option to tell the compiler it is actually a C source is not universally supported (this broke the test with clang-cl at least)
Martin Kroeker [Fri, 20 Mar 2020 22:00:06 +0000 (23:00 +0100)]
Merge pull request #2518 from shengyang-3390/dev
add ctest for srotm and modified ctest for srot.
Martin Kroeker [Fri, 20 Mar 2020 21:57:44 +0000 (22:57 +0100)]
Merge pull request #2519 from martin-frbg/issue2472
Fix cmake compilation with ifort on Windows
wjc404 [Fri, 20 Mar 2020 21:46:18 +0000 (21:46 +0000)]
Update param.h
wjc404 [Fri, 20 Mar 2020 21:42:10 +0000 (05:42 +0800)]
Add files via upload
Martin Kroeker [Fri, 20 Mar 2020 00:08:10 +0000 (01:08 +0100)]
Make ifort on Windows create lowercase symbols with appended underscore
tentative fix for #2472
Martin Kroeker [Fri, 20 Mar 2020 00:05:22 +0000 (01:05 +0100)]
Merge pull request #38 from xianyi/develop
rebase
shengyang [Wed, 18 Mar 2020 06:17:32 +0000 (14:17 +0800)]
add ctest for srotm and modified ctest for srot.
make sure that test cases cover all code path when kernel uses looping unrolling.
Martin Kroeker [Tue, 17 Mar 2020 09:12:53 +0000 (10:12 +0100)]
Merge pull request #2517 from wjc404/develop
Temporary fix for SKX STRSM
wjc404 [Tue, 17 Mar 2020 04:52:55 +0000 (12:52 +0800)]
Update KERNEL.SKYLAKEX
Martin Kroeker [Mon, 16 Mar 2020 20:59:55 +0000 (21:59 +0100)]
Merge pull request #2515 from zelong-1024/develop
[OpenBLAS]: benchmark for her/her2 LEVEL2 functions
Martin Kroeker [Mon, 16 Mar 2020 20:58:55 +0000 (21:58 +0100)]
Merge pull request #2513 from aaawuanjun/develop
[OpenBlas]: Add benchmark tpsv file and modify benchmark/Makefile
Martin Kroeker [Mon, 16 Mar 2020 20:58:34 +0000 (21:58 +0100)]
Merge pull request #2516 from wjc404/develop
AVX2 STRSM kernels
wjc404 [Mon, 16 Mar 2020 16:39:37 +0000 (16:39 +0000)]
Update KERNEL.ZEN
wjc404 [Mon, 16 Mar 2020 16:34:08 +0000 (00:34 +0800)]
AVX2 STRSM kernel
l00536773 [Mon, 16 Mar 2020 03:19:05 +0000 (11:19 +0800)]
[OpenBLAS]: benchmark for her/her2 LEVEL2 functions
[description]: benchmark for her/her2
[solution]: added benchmark for her/her2, modified makefile in benchmark
[dts]:
Martin Kroeker [Sat, 14 Mar 2020 13:21:30 +0000 (14:21 +0100)]
Merge pull request #2508 from liujingjue/develop
[OpenBLAS]:fix the iamax benchmark error
Martin Kroeker [Sat, 14 Mar 2020 12:27:40 +0000 (13:27 +0100)]
Merge pull request #2512 from martin-frbg/lapackh
Move declarations of lapack_complex_custom types outside the extern C
Martin Kroeker [Sat, 14 Mar 2020 12:08:36 +0000 (13:08 +0100)]
Merge pull request #2506 from xiaofengF/develop
Add benchmark for SPMV and fix segmentation fault when data size >= 50000
wuanjun 00447568 [Sat, 14 Mar 2020 01:11:08 +0000 (09:11 +0800)]
[OpenBlas]: Add benchmark tpsv file and modify benchmark/Makefile
[Description]: Solve lack of tpsv benchmark.
Martin Kroeker [Fri, 13 Mar 2020 22:16:35 +0000 (23:16 +0100)]
Merge pull request #2505 from aaawuanjun/develop
[OpenBlas]:Add benchmark tpmv.c and modify benchmark/Makefile